Compare commits

...

47 Commits

Author SHA1 Message Date
Feng Ruohang bf053386a4 docs: link compression and federation to maintained design records
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 18:26:03 +08:00
Feng Ruohang f99ed829b5 Merge pull request #213 from pgsty/codex/pr198-release-compat
fix: preserve multipart defaults and rollback compatibility
2026-09-16 16:42:56 +08:00
Feng Ruohang 956a1a8e33 Merge pull request #212 from pgsty/codex/docs-migration-cleanup
docs: consolidate repository entry points and retire migrated work records
2026-09-16 16:39:40 +08:00
Feng Ruohang 82f0a9828e fix: preserve multipart defaults and rollback compatibility
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 16:30:43 +08:00
Feng Ruohang 584b5a3a8f docs: link the canonical recovery runbooks
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 16:19:56 +08:00
Feng Ruohang 39f49d548f Merge remote-tracking branch 'origin/main' into codex/docs-migration-cleanup
Signed-off-by: Feng Ruohang <rh@vonng.com>

# Conflicts:
#	CONTRIBUTORS.md
2026-09-16 16:13:31 +08:00
Feng Ruohang a168576adb Merge pull request #208 from pgsty/codex/conditional-put-release-notes
docs: describe conditional PUT behavior and tested recovery guidance
2026-09-16 16:09:22 +08:00
Feng Ruohang e85c4c8dfb docs: preserve reviewed contributor records and daily card references
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 16:09:00 +08:00
Feng Ruohang 0e3c43778e Merge pull request #211 from pgsty/codex/daily-repository-cards
ci: update README repository cards daily at 00:00 UTC
2026-09-16 15:40:09 +08:00
Feng Ruohang 9101fe78db ci: refresh README repository cards daily at 00:00 UTC
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 15:27:48 +08:00
Feng Ruohang 82948d6306 Merge pull request #198 from mrjavadseydi/fix/issue-79-list-multipart-uploads
Retain the original contributor commit and integrate the reviewed durable listing, cancellation confirmation and migration fixes. Capacity acceptance and delayed creation-write fencing remain open in issue #79.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 13:40:10 +08:00
Feng Ruohang dfb4b2a1d2 fix: return HTTP 503 when multipart scan capacity is exhausted
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 13:24:52 +08:00
Feng Ruohang 143f6970d8 fix: make multipart discovery and cancellation explicit and bounded
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 13:16:37 +08:00
Feng Ruohang ea5dac5a99 Merge current main into multipart listing contribution
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 12:51:20 +08:00
Feng Ruohang 3c26a8b0b5 Merge pull request #209 from pgsty/codex/credit-console-share-reporter
fix: integrate the bounded Console sharing proxy and credit its reporter
2026-09-16 12:19:29 +08:00
Feng Ruohang 2fabd436c0 fix: select the bounded Console sharing proxy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 12:05:00 +08:00
Feng Ruohang 07c68d6054 docs: credit Jiri Pejchal for the Console sharing report
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 11:48:38 +08:00
Feng Ruohang 15090dc4fd docs: retire migrated working material and clarify documentation ownership
Remove migrated investigation and security documents, repair their incoming links, and align the English/Chinese READMEs with current module and release boundaries. Track AGENTS.md as the shared repository guide and make CLAUDE.md import it; keep working artifacts outside the repository.

Publish pgsty/silo.pgsty.com commit c7682185e2832b981629e1c17aef74fc778c6ca1 before publishing this cleanup: the new documentation routes are not live yet.

Validation: make rebrand-guard; 92 Markdown files checked for deletion-induced broken relative paths; 199 source-to-site references and anchors resolve in the paired site build; git diff --check.
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 11:22:48 +08:00
Feng Ruohang e791640dac docs: describe merged conditional PUT behavior and recovery guidance
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 11:21:16 +08:00
Feng Ruohang 9b4ae82a29 Merge pull request #207 from pgsty/codex/conditional-put-pools
fix(pools): evaluate cross-pool PUT conditions against the current object
2026-09-16 11:17:02 +08:00
Feng Ruohang 4620be394b test: synchronize conditional PUT disk fixtures
Replace unsynchronized getDisks swaps with backing disk-list updates under
erasureDisksMu, matching the existing GetDisks reader lock. Apply the same
helper to capacity and read-fault adapters while preserving nested restore
ordering.

Add a regression that overlaps fixture changes with the real IAM Walk
reader, and run the conditional PUT suite under the race detector in CI.
The regression reproduces the old fixture race; ten fixed race iterations
pass without warnings. Production conditional PUT behavior is unchanged.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 10:53:24 +08:00
Feng Ruohang 5e7d603083 fix(pools): evaluate cross-pool PUT conditions against the current object
Multi-pool PUT selected a destination by capacity and evaluated
If-Match/If-None-Match only against that destination's local object
state. An empty or stale destination could accept a stale ETag or
If-None-Match:* while another pool held the current object, replacing
it; a current ETag could instead be rejected with 412 or 404.

Under PUT's existing pools-layer object lock, resolve the comparison
object with objectPoolInfos (including draining pools), treat a latest
delete marker as absence, fail closed on unreadable pool metadata, and
clear an accepted callback before destination dispatch. Replica and
data-movement callbacks keep their addressed-version semantics and
metadata reconciliation.

Reproduced on 40220bd836 and RELEASE.2026-09-03T13-18-01Z with six
signed HTTP scenarios: four defect cases failed, two controls passed.

Refs #199

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:21:21 +08:00
Feng Ruohang fb7c406ddc Merge pull request #206 from pgsty/codex/main-consolidation-20260916
test: restore valid credentials in integration fixtures
2026-09-16 08:18:56 +08:00
Feng Ruohang 416826f61f Merge pull request #205 from pgsty/codex/release-notes-followups
docs: complete September correctness and upgrade notes
2026-09-16 08:16:23 +08:00
Feng Ruohang a2e2f3ee82 test: restore valid credentials in integration fixtures
Port the shell-fixture changes from ebc9937d97b27871dcc4bb91d4b5771d3550b76a. Match the existing minimum secret length consistently across startup, aliases and helper commands.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:08:08 +08:00
Feng Ruohang 70c7ec4a9f docs: link reviewed runbook sources before website deployment
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 08:04:53 +08:00
Feng Ruohang 2b722b7e92 docs: complete September correctness and upgrade notes
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 07:51:11 +08:00
mr javad seydi 4cbb074ccd fix: make multipart upload listing S3-compatible
Signed-off-by: mr javad seydi <seydi.birjand@gmail.com>
2026-09-15 23:17:55 +03:30
Feng Ruohang 40220bd836 Merge pull request #196 from pgsty/codex/merge-r4-r8
Complete R5, R6 and R8 on the main baseline containing R4 and R7.

Preserve the individual signed repairs, independent Opus 5 Max review, and integration validation evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 01:14:06 +08:00
Feng Ruohang df0dfa0a34 docs: record R4-R8 integration review and validation
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:58:31 +08:00
Feng Ruohang 80684fed59 chore: align integrated repair tests with contribution checks
Use the actual author and AGPL-3.0-or-later notices for new R6/R8 tests, preserving all test bodies and the Linux build tag. Record the existing replication ARN prefix used by the new R6 fixture in the compatibility inventory; no wire behavior changes.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:44:34 +08:00
Feng Ruohang 055030ea53 fix(http): honor configured request header deadlines
Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 0d48d32d7e038ae1ea5966f3d7e0cb86780a6311)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang aea3882c95 docs: record R6 integration verification
(cherry picked from commit d38edb2c46182d3a8fa96e040604493d20a4b478)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:56 +08:00
Feng Ruohang 0c61128d23 fix(replication): retry marker purges through persisted MRF
(cherry picked from commit cf381a7151ef25fc95ace5fedcd767fa19410de2)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 680eac66e4 fix(replication): preserve ordered tag deletions
Persist, send and reconcile empty tag states together with their revision
across COPY, PUT and multipart replication. Advance local tag mutations
under the existing locks and preserve current tags during replication ACK.

Cover signed HTTP, persistent single/multiple pool state, KMS, SSE-C key
rotation, ordering, retry and duplicate requests. Record real Opus plan
consensus, implementation review and local verification evidence.

Signed-off-by: Feng Ruohang <rh@vonng.com>
(cherry picked from commit 115fe8b12329d147adbaf817faa1737392ecbf9b)
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:37:55 +08:00
Feng Ruohang 9f3037e941 Merge pull request #194 from pgsty/codex/r7-replication-content-encoding
fix(replication): preserve normalized replica metadata
2026-09-16 00:31:03 +08:00
Feng Ruohang 5f00f6762f Merge main after R4 validation into R7 candidate
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:21:21 +08:00
Feng Ruohang af2b1794d3 Merge pull request #193 from pgsty/codex/r4-kms-tag-timestamp
fix(replication): preserve SSE-KMS tag timestamps
2026-09-16 00:20:10 +08:00
Feng Ruohang d371f77dcc docs: record final Opus 5 Max R7 implementation review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:10:31 +08:00
Feng Ruohang 022722a7a7 chore: align R4 contribution notices and record final review
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:08:09 +08:00
Feng Ruohang 03027727d1 fix(replication): preserve SSE-KMS tag timestamps
Carry the already-parsed source tagging timestamp through the KMS options
constructor so replica COPY can apply newer tag updates on explicitly or
automatically encrypted destinations.

Cover all option encryption modes and signed COPY persistence for newer,
stale, duplicate and timestamp-less updates, including bucket defaults.
Preserve the real Opus 5.0/max plan review, consensus and local validation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:03:52 +08:00
Feng Ruohang 4fcdf37ce6 fix(replication): preserve normalized replica metadata
Restore only the six replication-specific metadata fields after trust
validation, so streaming uploads retain their actual content encoding and
Snowball entries do not inherit ordinary metadata from the outer archive.

Include helper, authenticated PUT/COPY/multipart and Snowball regressions,
plus the R7 investigation, actual Opus 5 consensus and local verification.
The production change is based on PR #187 by Mikhail Khadarenka.

Co-authored-by: Mikhail Khadarenka <chodorenko@gmail.com>
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-16 00:00:17 +08:00
Feng Ruohang 9ebe81c1b3 Merge pull request #192 from pgsty/codex/iam-revision-tombstones
fix(iam): retain revocation versions through replay and recovery
2026-09-15 23:14:22 +08:00
Feng Ruohang e5f5c9e7f6 Merge pull request #191 from pgsty/codex/iam-peer-delete-reload
fix(iam): reload committed state on peer deletion notifications
2026-09-15 23:10:38 +08:00
Feng Ruohang 7b4cacc392 chore: align IAM file notices with contribution policy
Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang a0dd7dae9b fix(iam): resume site healing after leadership changes
Keep one healing loop per process across replication configuration reloads. Reacquire leadership after a lease is canceled and allow shutdown while waiting, so temporary quorum loss cannot permanently stop revocation propagation.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
Feng Ruohang 709d50a916 fix(iam): persist revocations across site replay and recovery
Retain source-ordered tombstones and parent grant boundaries across both IAM backends, cache reloads, and deliberate identity recreation. Reconcile deletions through a versioned, bounded replication protocol with restart-aware acknowledgements.

Cover inherited group grants, STS retention, same-key service recreation, absolute expiration, and failures after the durable commit. Document coordinated upgrades and the remaining consistency boundaries.

Signed-off-by: Feng Ruohang <rh@vonng.com>
2026-09-15 22:59:04 +08:00
197 changed files with 12711 additions and 16458 deletions
+6
View File
@@ -90,6 +90,12 @@ jobs:
- name: Run S3 Select tests under race detector
run: go test -race ./internal/s3select/... -count=1
- name: Run conditional PUT tests under race detector
run: go test -race ./cmd -run '^Test(PoolsConditionalPut|SinglePoolConditionalPutHTTP)' -count=1 -timeout=5m
- name: Run multipart listing and cancellation tests under race detector
run: go test -race ./cmd -run '^Test(MultipartListing|MultipartAbort|PaginateMultipartUploads|ListMultipartUploads)' -count=1 -timeout=5m
crosscompile:
name: Cross Compile
runs-on: ubuntu-latest
+100
View File
@@ -0,0 +1,100 @@
name: Repository Cards
on:
schedule:
# 00:00 UTC = 08:00 Asia/Shanghai. GitHub may queue scheduled runs.
- cron: "0 0 * * *"
workflow_dispatch:
push:
branches: [main]
paths:
- .github/workflows/repository-cards.yml
- .github/silo.svg
- buildscripts/repository-cards/**
pull_request:
branches: [main]
paths:
- .github/workflows/repository-cards.yml
- .github/silo.svg
- buildscripts/repository-cards/**
permissions:
contents: read
concurrency:
group: repository-cards-${{ github.event.pull_request.number || 'publish' }}
cancel-in-progress: false
jobs:
check:
name: Validate repository cards
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.13"
- run: python -m pip install -r buildscripts/repository-cards/requirements.txt
- run: python -m unittest discover -s buildscripts/repository-cards -p 'test_*.py' -v
publish:
name: Update README images
needs: check
if: github.repository == 'pgsty/silo' && github.ref == 'refs/heads/main' && github.event_name != 'pull_request'
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write
issues: read
pull-requests: read
env:
OUTPUT_BRANCH: codex/repository-cards
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.13"
- run: python -m pip install -r buildscripts/repository-cards/requirements.txt
- name: Load the generated-assets branch
shell: bash
run: |
set -euo pipefail
artifacts="$RUNNER_TEMP/repository-cards"
if git ls-remote --exit-code --heads origin "$OUTPUT_BRANCH"; then
git fetch --depth=1 origin "$OUTPUT_BRANCH"
git worktree add --detach "$artifacts" FETCH_HEAD
else
status=$?
# Exit 2 means no matching ref; transport/auth failures must stop.
if [ "$status" -ne 2 ]; then exit "$status"; fi
git worktree add --detach "$artifacts" HEAD
git -C "$artifacts" checkout --orphan "$OUTPUT_BRANCH"
git -C "$artifacts" rm -rf .
fi
- name: Refresh contributor and star cards
env:
GH_TOKEN: ${{ github.token }}
run: python buildscripts/repository-cards/update.py --output "$RUNNER_TEMP/repository-cards"
- name: Publish changed assets
shell: bash
run: |
set -euo pipefail
cd "$RUNNER_TEMP/repository-cards"
git add README.md history.json curated.json contributors.json \
contributors-light.svg contributors-dark.svg \
star-history-light.svg star-history-dark.svg
if git diff --cached --quiet; then
echo "Repository cards are already current."
exit 0
fi
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
git commit -s -m "chore: update repository cards $(date -u +%F)"
# A normal push preserves history and refuses concurrent overwrites.
git push origin "HEAD:refs/heads/$OUTPUT_BRANCH"
+2 -2
View File
@@ -62,10 +62,10 @@ dist/
.claude/
.codex/
AGENTS.md
CLAUDE.md
_bmad/
_bmad-output/
# Investigation and editorial working files belong outside this repository.
docs/investigations/
docs/security/
docs/rebranding.md
.release-sign/
+84
View File
@@ -0,0 +1,84 @@
# SILO Repository Guide
## Project map
- Server: `pgsty/silo` (this checkout, default branch `main`)
- Console: `pgsty/silo-console` (usual sibling checkout `../silo-console`)
- Client: `pgsty/mc`, shipped as `mcli` (usual sibling checkout `../mc`)
- Shared packages: `pgsty/silo-pkg` (usual sibling checkout `../silo-pkg`)
- Documentation: `pgsty/silo.pgsty.com` (usual sibling checkout `../silo.pgsty.com`),
published at <https://silo.pgsty.com/>
The server executable, package, systemd service, and container image use `silo`;
the public Docker image is `docker.io/pgsty/silo`. The server module remains
`github.com/minio/minio`. Consult the current `go.mod` and release configuration
for selected component versions and compatibility replacements.
## Supported stack and compatibility policy
The maintained, release-gating product graph is the coordinated PGSTY stack:
`silo` + `silo-console` + `mc` + `silo-pkg`.
Keep compatibility with upstream MinIO/MC on a best-effort basis. Preserve
inexpensive wire, configuration, CLI, migration, and import compatibility when
it helps users, and document known differences. Do not infer from source
lineage, MinIO-compatible protocols, retained `MINIO_*`/`MC_*` names, or an
inherited upstream test that unmodified upstream MinIO is a supported release
target.
Upstream-only compatibility checks may remain as advisory evidence, but they
must not force a downgrade of a maintained SILO component, a fork-only API
workaround, or a release block. Make upstream compatibility a hard gate only
when the user explicitly requests that scope.
## Dependency selection
- When PGSTY maintains a component, import and require it directly under its own
module path. In particular, prefer `github.com/pgsty/silo-pkg/v3` over
`github.com/minio/pkg/v3` in maintained SILO source.
- Do not use an upstream package merely to make unmodified upstream MinIO or MC
compile. Unavoidable transitive upstream modules should be documented and
kept separate from the maintained product dependency.
- `github.com/minio/minio-go/v7` is the explicit exception: use the verified
upstream module/commit while it contains the required fixes; do not recreate
a SILO fork without a concrete functional divergence.
- Check the current `go.mod` before changing versions. Coordinate breaking
import-path changes across package, client, Console, and server releases.
## Documentation and delivery
The companion site owns product, operations, migration, security-advisory,
design, and release documentation. Maintain English and Chinese content in
that repository and link to its canonical URLs from this one. The site's
`content/docs/` is the documentation entry point; detailed pages live under
`content/operations/`, `administration/`, `reference/`, `compatibility/`,
`about/`, and `blog/`, following the existing structure.
Keep repository entry points and contributor instructions here. Retained
upstream documents, examples, and test fixtures under `docs/` may be updated
when a code or tooling change requires it; do not build a second documentation
site in this tree.
Investigation plans, prompts, AI session transcripts, execution logs, temporary
reports, and private security material belong outside the checkout, for example
in a task-specific directory under `~/tmp/`. Do not recreate
`docs/investigations/`, `docs/security/`, or `docs/rebranding.md`, or move their
working material to another repository directory. Extract reusable, verified
knowledge into the companion site; keep private evidence outside public Git.
Compatibility pages may continue to describe similarities with upstream, but
must label that compatibility as best effort. The supported and tested server
for Console and mcli administration is `pgsty/silo`.
Before removing or moving documentation, check repository links and scripts,
the companion site's links and anchors, and references from this guide. Use
fixed commit URLs for historical evidence that should survive deletion from
the current tree. Run the site's `make check` for site changes and this
repository's `make rebrand-guard` for documentation cleanup that changes the
identifier inventory. When new public URLs are involved, publish the site
before publishing source changes that depend on those URLs.
Keep changes in the repository that owns the affected surface. Treat every
repository commit, tag, release, image, documentation update, and deployment as
a separate deliverable. Follow `CONTRIBUTING.md`, including DCO sign-off
(`git commit -s`). `CLAUDE.md` imports this guide so both agents use one policy.
+101 -6
View File
@@ -9,9 +9,36 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
### Authorization and security
- Restrict embedded Console's anonymous sharing proxy to object-content GETs
at the configured S3 origin, and reject every redirect. Internal metrics,
system paths and non-download S3 operations cannot be reached through it.
Normal public, presigned and versioned downloads remain available without a
new setting; a full sharing-disable switch is not introduced. See
[Console #56](https://github.com/pgsty/silo-console/pull/56) and the
[design record](https://github.com/pgsty/silo-console/issues/52).
Thanks to Jiri Pejchal (@jiri-pejchal) for the report.
- Persist IAM deletion revisions and parent revocation boundaries so stale site
events cannot restore deleted identities, policies or their older grants
(#191, #192). Peer deletion notifications reload committed storage; deliberate
recreation requires a newer revision, and credentials issued before the
parent's revocation remain invalid.
**Coordinated upgrade required:** upgrade every participating node and site.
Mixed old/new nodes sharing an IAM backend and rolling downgrade are
unsupported. Back up complete IAM storage and encryption material; an admin
export of live records omits deletion history. Reissue credentials for
recreated parents and explicitly reconcile pre-upgrade revocations whose
history is already lost. Restoring an older backup can lose later revocations;
keep affected sites isolated until reconciliation/rekeying is complete. See
[the operator runbook](https://silo.pgsty.com/operations/replication/iam-upgrade/).
- Enforce an absolute HTTP/1 request-header deadline through the connection
wrapper (#196). Repeated small reads no longer extend that deadline, and
`--read-header-timeout` / `MINIO_READ_HEADER_TIMEOUT` now reaches the HTTP
server. HTTP/1 request bodies retain the rolling idle timeout; this does not
impose a total upload/download duration. A shorter setting also constrains
TLS handshake reads. The wrapper's strict header mode is not applied to HTTP/2.
- Reject unsigned `x-amz-*` request headers that could turn a signed PUT into a
copy of another object accessible to the signer (SN-2026-011). The latest
public Server is affected; the fix is on main. See [the advisory ledger](docs/security/advisories.md).
public Server is affected; the fix is on main. See [the advisory ledger](https://silo.pgsty.com/about/security-advisories/).
- Align signed request fields with policy conditions and enforce header-only
presigned payload checksums. See [the signed-header review](https://silo.pgsty.com/blog/design/signed-header-coverage/).
- **Breaking policy semantics:** separate self-service `admin:ChangeMyPassword`
@@ -22,17 +49,85 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
### Object storage and replication
- Make `ListMultipartUploads` discover quorum-valid uploads from durable state
across pools, erasure sets and drives, then apply S3 prefix, delimiter,
marker, ordering and 1,000-entry pagination semantics globally (#198). New uploads
store their canonical bucket and key as reserved fields in the existing
quorum-written `xl.meta`; completion removes those upload-only fields. Native
markers remain usable after their upload is completed or canceled. Strict
listing returns a diagnostic 503 for legacy uploads or uncertain coverage;
the default remains the released exact-key/cache-based `legacy` behavior.
Opt into strict mode only through `MINIO_API_MULTIPART_LISTING=strict`, after
upgrading every writer, draining old uploads, checking the read-only admin
`multipart-preflight` report and validating scan capacity. The process-only
setting is not persisted into shared API configuration. Historical
`multipart_listing` keys are ignored and can be removed with a targeted
`mcli admin config reset ALIAS api multipart_listing` before rollback. Per-process
admission, directory-entry, worker and time budgets bound scan scheduling;
each page still scans durable state. See [issue #79](https://github.com/pgsty/silo/issues/79)
and its [design record](https://silo.pgsty.com/blog/design/list-multipart-uploads/).
Thanks to mr javad seydi (@mrjavadseydi) for the original implementation.
- Retain released read-quorum and best-effort multipart cancellation in default
legacy mode. Strict mode requires majority deletion acknowledgements and
permits retries below read quorum. When most drives were already empty,
failed deletion of an observed remnant now returns 503 instead of being
masked by empty-drive successes. Wrong-key or wrong-bucket cancellation
preserves the valid upload's cache entry; successful cancellation notifies
peers with the request context after the distributed lock is released.
The HTTP response for an absent
upload remains 204; this is not proof of physical cleanup. **Known boundary:**
delayed creation writes can still restore an upload after cancellation;
this change does not add a durable creation fence.
- Preserve object tags during multi-pool metadata reconciliation by reading the
resolved tag field together with its revision (#189). Previously, reconciliation
could replace existing tags with an empty value.
- Preserve the tag revision on SSE-KMS metadata replication (#193), and advance
tag revisions monotonically on local PUT/DELETE tagging (#196). Empty tags
participate in reconciliation as an ordered deletion, preventing older
events from restoring removed tags. SSE-C key rotation also retains the tag
revision. Malformed historical revisions can fail and retry; their missing
history is not reconstructed by the upgrade.
- Complete delete-marker version purges and preserve their identity and retry
state through MRF recovery (#196). Recovery accepts a 405 marker response only
when its version, bucket, object name and modification time match the task.
Purge audit status is normalized from `COMPLETE` to `COMPLETED`.
Thanks to Julien Laurenceau (@julienlau) for the investigation and proposed
fix in #184 that helped shape this follow-up.
- Restore only the six replication-specific metadata fields after ordinary
request metadata extraction (#194). This prevents transport-only `aws-chunked`
from being stored as Content-Encoding while preserving the signed-header
protections. Trusted Snowball entries no longer inherit the outer archive's
ordinary metadata. Thanks to Mikhail Khadarenka (@chodorenko) for the fix in #187.
**Existing data:** these repairs prevent new errors; they do not scan or rewrite
historical object metadata, recover lost tags or prove that old purge work has
converged. Follow the [read-only audit procedure](https://silo.pgsty.com/operations/replication/replica-metadata-audit/)
before planning any repair of stored state.
- Evaluate conditional multipart completion against the logical current object
across all pools while holding the existing object lock. A stale `If-Match`
can no longer replace newer data in another pool, and the current ETag is no
longer rejected because the upload resides next to an older copy. Conditions
are evaluated once; a current delete marker counts as an absent object.
**Availability change:** if metadata cannot be read from any pool, conditional
**Availability change:** if any pool's metadata cannot be read, conditional
completion fails even when another pool can still serve GET/HEAD. This also
applies when the unreadable pool may not hold the object: absence cannot be
verified. Retry after the pool recovers. Unconditional completion and the
single-pool path retain their existing behavior.
- Evaluate ordinary multi-pool conditional PUT against the logical current
object across all pools, including draining pools, under the existing object
lock (#207). A stale destination copy no longer accepts a stale ETag or rejects
the current one; a current delete marker is treated as absence.
**Availability change:** if any pool's object metadata cannot be verified,
the condition fails even when GET can use another pool; read-quorum failures
return 503. Restore readability or heal before retrying. Unconditional PUT,
single-pool conditions and internal replication retain their existing behavior.
A public condition with a destination `versionId` compares the current object
while preserving the requested write version. This change does not retire
stale copies in other pools, undo historical accepted overwrites or provide
a new global clock-ordering guarantee. The multipart-completion repair in #190
neither introduced nor repaired this separate PUT defect.
- Reconcile ordinary single-object version DELETE across all pools, including
null versions, delete markers and unqualified directory-marker DELETE. This
applies the deletion to every resolved pool copy under existing quorum
@@ -50,7 +145,7 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
its tracker, mover, scanner hooks, configuration, XML actions and metrics.
Accept and ignore retired configuration/XML and preserve ordinary statistics
when reading v9 caches. See [migration notes](docs/bucket/lifecycle/access-tiering-removal.md).
The [decision record](docs/investigations/access-tiering-revert.md) preserves
The [decision record](https://silo.pgsty.com/compatibility/access-tiering-removal/) preserves
the feature's introduction, subsequent fixes, rollback scope and review history.
- Preserve the independent multi-pool write, metadata, healing and conditional
deletion fixes from PR #178, including shared remote-tier reference protection.
@@ -61,10 +156,10 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
- Repair federated CopyObject checksums, destination timestamps, reserved
metadata, encrypted-object forwarding, legal hold and KMS context.
- Make resync counters, target selection, cancellation and worker lifetimes
reflect actual work; complete delete-marker purges and report bounded MRF drops.
reflect actual work, and report bounded MRF drops.
- Converge bucket metadata with deterministic source state, deletion tombstones,
creation time recovery and diagnostics. The mixed-version export gate requires
coordinated upgrades before tombstones are exported. See [the #77 record](docs/investigations/issue-77-current.md).
coordinated upgrades before tombstones are exported. See [the #77 record](https://silo.pgsty.com/blog/design/bucket-metadata-convergence/).
- Include per-bucket CORS in metadata export/import, close metadata publication
and logger races, and report effective bucket quotas in metrics.
@@ -73,7 +168,7 @@ and [complete commit range](https://github.com/pgsty/silo/compare/RELEASE.2026-0
- Restore embedded Console login over loopback TLS, trusted-proxy handling and
all four WebSocket connection limits. Preserve Go TLS defaults across transports.
- Directly require `github.com/pgsty/silo-pkg/v3` v3.14.0; select Console
`v0.0.0-20260913015128-417559bb2c97` and MC
`v0.0.0-20260916034812-56dfe455ac2f` and MC
`v0.0.0-20260913012246-4f609a4da3bb` with explicit PGSTY replacements.
- Pin upstream minio-go `v7.3.1-0.20260910142817-60bd07042d49`; refresh Go x/*
modules and security fixes including bounded AMQP frame handling. Keep Go
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+10
View File
@@ -77,6 +77,16 @@ test evidence, compatibility notes, and documentation impact. Public product
documentation is owned by the separate
[`pgsty/silo.pgsty.com`](https://github.com/pgsty/silo.pgsty.com) repository.
## Contributor recognition
Every human issue or pull-request author is recognized in [CONTRIBUTORS.md](CONTRIBUTORS.md),
including open issues, draft PRs, and PRs closed without merging. Merged fixes,
adopted proposals, and reports that lead to fixes receive greater prominence;
first participation guides the remaining order. Work incorporated through a
later PR retains credit without changing the original PR's recorded status.
Security disclosures are credited with the reporter's agreement. DCO sign-off
applies to code commits, not to opening an issue.
## Licensing of Contributions
Code contributions to PGSTY SILO (`pgsty/silo`) are accepted under the
+116 -120
View File
File diff suppressed because one or more lines are too long
+33 -49
View File
@@ -38,7 +38,7 @@
## Current release and main branch
The latest published Server is [20260903](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-09-03T13-18-01Z).
As of 2026-09-13, the main branch has newer security, storage, Console and
As of 2026-09-16, the main branch has newer security, storage, Console and
shared-package changes that have not shipped in a Server release. See
[CHANGELOG.md](CHANGELOG.md) and the [component version matrix](https://silo.pgsty.com/compatibility/versions/)
for the exact release/source boundary, including SN-2026-011 and password-policy migration.
@@ -93,13 +93,16 @@ Every release ships checksums, SPDX SBOMs, Sigstore-signed manifests, and GitHub
## Compatibility
The S3 API, `MINIO_*` variables, `minio_*` metrics, `x-minio-*` headers, `/minio/*` routes, the `github.com/minio/*` import paths, and the on-disk format (including `.minio.sys`) are preserved and held in place by a CI compatibility check. Only Silo-owned delivery surfaces change: the `silo` executable, package, service, Helm chart, and container image — no `minio` binary alias is installed.
Silo preserves S3 and storage-format compatibility, including existing `MINIO_*` variables, `minio_*` metrics, `x-minio-*` headers, `/minio/*` routes, and `.minio.sys` data. CI guards selected compatibility identifiers; release notes document intentional security and behavior changes. Silo-owned delivery surfaces use the `silo` executable, package, service, Helm chart, and container image; no `minio` server binary alias is installed.
The supported release stack is `pgsty/silo` + `pgsty/silo-console` + `pgsty/mc` + `pgsty/silo-pkg`; compatibility with unmodified upstream MinIO/MC is best effort. The Server, Console, and client retain their historical module paths where needed, while maintained code imports `github.com/pgsty/silo-pkg/v3` directly. The SDK `github.com/minio/minio-go/v7` is an explicit upstream dependency. See the current [go.mod](go.mod) for versions and replacements.
Every divergence from upstream is listed in the code-verified [compatibility audit](https://silo.pgsty.com/compatibility/server/). Treat each release as a downstream upgrade: pin versions, read the [release notes](https://silo.pgsty.com/tags/silo/), and keep a rollback path.
### TLS and Go upgrades
TLS key exchange follows Go's defaults across the S3 listener, node links,
The following TLS repair is on main and is not included in Server 20260903.
With that repair, TLS key exchange follows Go's defaults across the S3 listener, node links,
replication, identity providers, etcd, and external HTTP services. If an endpoint
cannot accept ML-KEM, `GODEBUG=tlsmlkem=0` disables the default hybrid exchanges
for the process; certificate verification remains enabled. This option does not
@@ -115,7 +118,16 @@ values to restore Keychain trust. Explicit certificates in the configured `CAs`
directory remain additive to the selected root pool.
Go 1.27 binaries require macOS 13 or later. See the
[Go release notes](https://go.dev/doc/go1.27) and the
[SILO stack investigation](docs/investigations/go127-stack.md).
[Go 1.27 TLS and OIDC discovery guide](https://silo.pgsty.com/blog/design/go127-tls-oidc-discovery/).
## Documentation ownership
User documentation is maintained at [silo.pgsty.com](https://silo.pgsty.com/docs/),
with source in [pgsty/silo.pgsty.com](https://github.com/pgsty/silo.pgsty.com).
The remaining `docs/` tree contains inherited references, examples, and tooling
fixtures. Investigation logs, AI work records, and temporary reports are kept
outside this repository; reusable findings belong in the companion site.
See [AGENTS.md](AGENTS.md) for repository ownership and maintenance rules.
## Security & Contributing
@@ -123,53 +135,25 @@ Report vulnerabilities privately as described in [`SECURITY.md`](SECURITY.md); e
## Contributors
**41 community contributors** build SILO, Console, mcli, shared packages, and related projects. The list includes maintainers and every human Issue or PR author, ordered by merged PRs, other PRs, then issue reports. Gold rings highlight significant contributions.
<!-- Generated by silo.pgsty.com/bin/contributors.py. -->
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — Maintains SILO, Console, mcli, shared packages, releases, and documentation"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — Implemented per-bucket CORS configuration and enforcement"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — Fixed effective bucket quota metrics; proposed access-frequency ILM"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — Added ChecksumType to multipart completion responses"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — Fixed missing-bucket ListObjects semantics"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — Repaired bucket notification streaming"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — Patched CVE-2026-34986 in go-jose"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — Patched CVE-2026-39883 in OpenTelemetry"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — Repointed documentation links to the SILO portal"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — Contributed the replicated SSE-C plaintext part-size fix"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — Reported and proposed explicit-version delete authorization"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — Reported and proposed conditional DELETE support for If-Match"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — Proposed the distroless container image and dependency automation"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — Proposed robustness and goroutine improvements"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — Proposed dependency updates"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — Reported multipart and streaming checksum defects and missing-bucket semantics"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — Reported LDAP TLS and Console login regressions"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — Reported the unsigned-header CopyObject cross-object read (SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — Reported bucket quota metrics reading a deprecated field"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — Reported missing user, group, and defaults in Debian packages"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — Reported missing RPM package signatures"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — Reported the missing client in release tarballs"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — Reported the missing client in the container image"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — Requested the migration guide from upstream MinIO"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — Reported NATS JWT credentials and target reload issues"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — Reported ListMultipartUploads prefix and pagination semantics"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — Reported indeterminate progress for prefix downloads"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — Reported Console lifecycle management and file preview gaps"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — Requested cpuv1 support and reported a workflow token failure"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — Reported inter-node I/O timeouts in ReadFileStreamHandler"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — Requested KES-compatible external KMS and OpenBao support"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — Reported the missing documentation navigation"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — Requested a maintained Helm chart"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — Reported missing Console tiering and site replication"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — Reported the unusable profiling option"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — Reported Chinese documentation availability"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — Requested Windows build guidance"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — Asked for clarity on Helm chart and operator options"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — Proposed making the SILO Operator easier to discover"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — Proposed CNCF Sandbox governance"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — Asked about community support and image maintenance"></a>
</p>
Every human issue or pull-request author is part of the SILO community, including open and unmerged work. Merged fixes, adopted proposals, and actionable reports receive priority, with first participation guiding the remaining order. Gold rings highlight reviewed significant contributions.
[View the full contribution record](CONTRIBUTORS.md) for each person's proposals, fixes, and reports.
<a href="CONTRIBUTORS.md">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/contributors-dark.svg">
<img src="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/contributors-light.svg" alt="SILO community contributors">
</picture>
</a>
[View contribution notes and actual PR status](CONTRIBUTORS.md).
## Star History
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/star-history-dark.svg">
<img src="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/star-history-light.svg" alt="SILO GitHub star history">
</picture>
## Background
+34 -47
View File
@@ -38,7 +38,7 @@
## 当前发行版与主分支
最新已发布的 Server 仍为 [20260903](https://github.com/pgsty/silo/releases/tag/RELEASE.2026-09-03T13-18-01Z)。
截至 2026-09-13,主分支已合入更新的安全、存储、Console 与共享包改动,但尚未发布新 Server。
截至 2026-09-16,主分支已合入更新的安全、存储、Console 与共享包改动,但尚未发布新 Server。
准确的已发布/源码边界见 [CHANGELOG.md](CHANGELOG.md) 与[组件版本矩阵](https://silo.pgsty.com/zh/compatibility/versions/),
其中包括 SN-2026-011 修复状态与密码权限迁移要求。
@@ -92,63 +92,50 @@ docker exec silo mcli mb local/demo && docker exec silo mcli ls local
## 兼容性
S3 API、`MINIO_*` 环境变量、`minio_*` 指标、`x-minio-*` 头、`/minio/*` 路由、`github.com/minio/*` 导入路径与磁盘格式(含 `.minio.sys`)原样保留,并由 CI 兼容性门禁冻结。只有 Silo 自有交付面改名:`silo` 可执行文件、软件包、服务、Helm Chart 与容器镜像 —— 原生交付物不会安装 `minio` 二进制别名。
Silo 保留 S3 与存储格式兼容性,包括既有 `MINIO_*` 环境变量、`minio_*` 指标、`x-minio-*` 头、`/minio/*` 路由与 `.minio.sys` 数据。CI 守卫检查选定的兼容性标识,有意的安全与行为变化在发布说明中记录。Silo 自有交付面使用 `silo` 可执行文件、软件包、服务、Helm Chart 与容器镜像;原生交付物不会安装 `minio` 服务端二进制别名。
正式支持和发布验收的组合为 `pgsty/silo` + `pgsty/silo-console` + `pgsty/mc` + `pgsty/silo-pkg`,对未修改的上游 MinIO/MC 尽最大努力保持兼容。Server、Console 与客户端按需保留历史模块路径,维护源码直接导入 `github.com/pgsty/silo-pkg/v3`;SDK `github.com/minio/minio-go/v7` 是明确保留的上游依赖。具体版本与 replace 以当前 [go.mod](go.mod) 为准。
与上游的全部分歧,以逐项核验代码的[兼容性审计](https://silo.pgsty.com/zh/compatibility/server/)形式维护。每个版本仍应视为下游升级:锁定版本,阅读[版本说明](https://silo.pgsty.com/zh/tags/silo/),并保留回滚路径。
### TLS 与 Go 升级
Server 恢复 Go 默认密钥交换策略的修复已在 main,尚未包含在 Server 20260903。
`GODEBUG=tlsmlkem=0`、`tlssecpmlkem=0` 的适用范围、macOS 根证书来源变化,
以及 OIDC discovery 的诊断方法见 [Go 1.27 TLS 与 OIDC 指南](https://silo.pgsty.com/zh/blog/design/go127-tls-oidc-discovery/)。
## 文档归属
用户文档统一维护在 [silo.pgsty.com](https://silo.pgsty.com/zh/docs/),源码位于
[pgsty/silo.pgsty.com](https://github.com/pgsty/silo.pgsty.com)。本仓库保留的 `docs/`
主要是继承的参考材料、示例和工具测试夹具。调查日志、AI 工作记录与临时报告放在仓库外;
可复用的结论应整理进伴生文档站。仓库职责与维护规则见 [AGENTS.md](AGENTS.md)。
## 安全与贡献
请按照 [`SECURITY.md`](SECURITY.md) 私密报告漏洞;每项修复都会发布公开[安全公告](https://silo.pgsty.com/zh/blog/security/)。本项目不要求签署 CLA:贡献按 AGPL-3.0-or-later(inbound=outbound)接收,只需 DCO 签署(`git commit -s`),详见 [`CONTRIBUTING.md`](CONTRIBUTING.md)。
## 贡献者
**41 位社区贡献者**共同建设 SILO、Console、mcli、公共包与相关项目。名单包含维护者,以及所有提出 Issue 或 PR 的真人作者;按已合并 PR、其他 PR、Issue 报告排序,黄圈标记显著贡献。
<!-- Generated by silo.pgsty.com/bin/contributors.py. -->
<p align="center">
<a href="https://github.com/Vonng"><img src="https://silo.pgsty.com/images/contributors/Vonng.svg" width="60" height="60" alt="@Vonng" title="@Vonng — 维护 SILO、Console、mcli、公共包、发行与文档"></a>
<a href="https://github.com/h5vx"><img src="https://silo.pgsty.com/images/contributors/h5vx.svg" width="60" height="60" alt="@h5vx" title="@h5vx — 实现单桶 CORS 配置与请求执行"></a>
<a href="https://github.com/mrjavadseydi"><img src="https://silo.pgsty.com/images/contributors/mrjavadseydi.svg" width="60" height="60" alt="@mrjavadseydi" title="@mrjavadseydi — 修复有效桶配额指标,并提交按访问频率分层的 ILM 方案"></a>
<a href="https://github.com/Dansyuqri"><img src="https://silo.pgsty.com/images/contributors/Dansyuqri.svg" width="60" height="60" alt="@Dansyuqri" title="@Dansyuqri — 为分片上传完成响应补充 ChecksumType"></a>
<a href="https://github.com/ycjlin"><img src="https://silo.pgsty.com/images/contributors/ycjlin.svg" width="60" height="60" alt="@ycjlin" title="@ycjlin — 修复缺失桶的 ListObjects 语义"></a>
<a href="https://github.com/pinginfo"><img src="https://silo.pgsty.com/images/contributors/pinginfo.svg" width="60" height="60" alt="@pinginfo" title="@pinginfo — 修复桶通知的流式输出"></a>
<a href="https://github.com/ZouhairCharef"><img src="https://silo.pgsty.com/images/contributors/ZouhairCharef.svg" width="60" height="60" alt="@ZouhairCharef" title="@ZouhairCharef — 修复 go-jose 中的 CVE-2026-34986"></a>
<a href="https://github.com/mfredenhagen"><img src="https://silo.pgsty.com/images/contributors/mfredenhagen.svg" width="60" height="60" alt="@mfredenhagen" title="@mfredenhagen — 修复 OpenTelemetry 中的 CVE-2026-39883"></a>
<a href="https://github.com/waterkip"><img src="https://silo.pgsty.com/images/contributors/waterkip.svg" width="60" height="60" alt="@waterkip" title="@waterkip — 将文档链接指向 SILO 门户"></a>
<a href="https://github.com/mikemikimike"><img src="https://silo.pgsty.com/images/contributors/mikemikimike.svg" width="60" height="60" alt="@mikemikimike" title="@mikemikimike — 提交 SSE-C 复制分片明文尺寸修复"></a>
<a href="https://github.com/metaneutrons"><img src="https://silo.pgsty.com/images/contributors/metaneutrons.svg" width="60" height="60" alt="@metaneutrons" title="@metaneutrons — 报告并提交显式版本删除鉴权方案"></a>
<a href="https://github.com/magicxor"><img src="https://silo.pgsty.com/images/contributors/magicxor.svg" width="60" height="60" alt="@magicxor" title="@magicxor — 报告并提交 DELETE If-Match 条件请求支持方案"></a>
<a href="https://github.com/davinkevin"><img src="https://silo.pgsty.com/images/contributors/davinkevin.svg" width="60" height="60" alt="@davinkevin" title="@davinkevin — 提交 distroless 容器镜像与依赖自动更新方案"></a>
<a href="https://github.com/lem21h"><img src="https://silo.pgsty.com/images/contributors/lem21h.svg" width="48" height="48" alt="@lem21h" title="@lem21h — 提交健壮性与 goroutine 改进"></a>
<a href="https://github.com/sulin37392"><img src="https://silo.pgsty.com/images/contributors/sulin37392.svg" width="48" height="48" alt="@sulin37392" title="@sulin37392 — 提交依赖更新"></a>
<a href="https://github.com/cbornet"><img src="https://silo.pgsty.com/images/contributors/cbornet.svg" width="60" height="60" alt="@cbornet" title="@cbornet — 报告分片与流式校验和缺陷及缺失桶语义问题"></a>
<a href="https://github.com/vampywiz17"><img src="https://silo.pgsty.com/images/contributors/vampywiz17.svg" width="60" height="60" alt="@vampywiz17" title="@vampywiz17 — 报告 LDAP TLS 与 Console 登录回归"></a>
<a href="https://github.com/orenyomtov"><img src="https://silo.pgsty.com/images/contributors/orenyomtov.svg" width="60" height="60" alt="@orenyomtov" title="@orenyomtov — 报告未签名头导致的 CopyObject 跨对象读取(SN-2026-011)"></a>
<a href="https://github.com/mumu-lab"><img src="https://silo.pgsty.com/images/contributors/mumu-lab.svg" width="48" height="48" alt="@mumu-lab" title="@mumu-lab — 报告桶配额指标读取已弃用字段的问题"></a>
<a href="https://github.com/jvasile"><img src="https://silo.pgsty.com/images/contributors/jvasile.svg" width="48" height="48" alt="@jvasile" title="@jvasile — 报告 Debian 包缺少用户、用户组与默认配置"></a>
<a href="https://github.com/pmezhuev"><img src="https://silo.pgsty.com/images/contributors/pmezhuev.svg" width="48" height="48" alt="@pmezhuev" title="@pmezhuev — 报告 RPM 包缺少 GPG 签名"></a>
<a href="https://github.com/TLINDEN"><img src="https://silo.pgsty.com/images/contributors/TLINDEN.svg" width="48" height="48" alt="@TLINDEN" title="@TLINDEN — 报告发布压缩包缺少客户端"></a>
<a href="https://github.com/makinikm"><img src="https://silo.pgsty.com/images/contributors/makinikm.svg" width="48" height="48" alt="@makinikm" title="@makinikm — 报告容器镜像缺少客户端"></a>
<a href="https://github.com/meesudzu"><img src="https://silo.pgsty.com/images/contributors/meesudzu.svg" width="48" height="48" alt="@meesudzu" title="@meesudzu — 提出从上游 MinIO 迁移的指南需求"></a>
<a href="https://github.com/kuldeep-link11"><img src="https://silo.pgsty.com/images/contributors/kuldeep-link11.svg" width="48" height="48" alt="@kuldeep-link11" title="@kuldeep-link11 — 报告 NATS JWT 凭据与通知目标重载问题"></a>
<a href="https://github.com/sargarass"><img src="https://silo.pgsty.com/images/contributors/sargarass.svg" width="48" height="48" alt="@sargarass" title="@sargarass — 报告 ListMultipartUploads 前缀与分页语义问题"></a>
<a href="https://github.com/liuhaodongliu990-cmyk"><img src="https://silo.pgsty.com/images/contributors/liuhaodongliu990-cmyk.svg" width="48" height="48" alt="@liuhaodongliu990-cmyk" title="@liuhaodongliu990-cmyk — 报告前缀下载进度显示异常"></a>
<a href="https://github.com/Xavier-777"><img src="https://silo.pgsty.com/images/contributors/Xavier-777.svg" width="48" height="48" alt="@Xavier-777" title="@Xavier-777 — 报告 Console 生命周期管理与文件预览缺失"></a>
<a href="https://github.com/spaceg00se-r"><img src="https://silo.pgsty.com/images/contributors/spaceg00se-r.svg" width="48" height="48" alt="@spaceg00se-r" title="@spaceg00se-r — 提出 cpuv1 支持需求并报告工作流令牌错误"></a>
<a href="https://github.com/kh0mka"><img src="https://silo.pgsty.com/images/contributors/kh0mka.svg" width="48" height="48" alt="@kh0mka" title="@kh0mka — 报告 ReadFileStreamHandler 节点间 I/O 超时"></a>
<a href="https://github.com/bagutzu"><img src="https://silo.pgsty.com/images/contributors/bagutzu.svg" width="48" height="48" alt="@bagutzu" title="@bagutzu — 提出兼容 KES 的外部 KMS 与 OpenBao 支持需求"></a>
<a href="https://github.com/DestroyLee"><img src="https://silo.pgsty.com/images/contributors/DestroyLee.svg" width="48" height="48" alt="@DestroyLee" title="@DestroyLee — 报告文档目录导航缺失"></a>
<a href="https://github.com/mosesdd"><img src="https://silo.pgsty.com/images/contributors/mosesdd.svg" width="48" height="48" alt="@mosesdd" title="@mosesdd — 提出维护 Helm Chart 的需求"></a>
<a href="https://github.com/zylpsrs"><img src="https://silo.pgsty.com/images/contributors/zylpsrs.svg" width="48" height="48" alt="@zylpsrs" title="@zylpsrs — 报告 Console 缺少分层与站点复制"></a>
<a href="https://github.com/heroes1412"><img src="https://silo.pgsty.com/images/contributors/heroes1412.svg" width="48" height="48" alt="@heroes1412" title="@heroes1412 — 报告性能分析选项不可用"></a>
<a href="https://github.com/redfoxfox"><img src="https://silo.pgsty.com/images/contributors/redfoxfox.svg" width="48" height="48" alt="@redfoxfox" title="@redfoxfox — 报告中文文档站点不可用"></a>
<a href="https://github.com/jiadzh"><img src="https://silo.pgsty.com/images/contributors/jiadzh.svg" width="48" height="48" alt="@jiadzh" title="@jiadzh — 提出 Windows 构建指导需求"></a>
<a href="https://github.com/AntonOfTheWoods"><img src="https://silo.pgsty.com/images/contributors/AntonOfTheWoods.svg" width="48" height="48" alt="@AntonOfTheWoods" title="@AntonOfTheWoods — 提出明确 Helm Chart 与 Operator 选项的需求"></a>
<a href="https://github.com/chalukyaj"><img src="https://silo.pgsty.com/images/contributors/chalukyaj.svg" width="48" height="48" alt="@chalukyaj" title="@chalukyaj — 提出改善 SILO Operator 可发现性的建议"></a>
<a href="https://github.com/nsanitate"><img src="https://silo.pgsty.com/images/contributors/nsanitate.svg" width="48" height="48" alt="@nsanitate" title="@nsanitate — 提出加入 CNCF Sandbox 的治理建议"></a>
<a href="https://github.com/Kesavaambati"><img src="https://silo.pgsty.com/images/contributors/Kesavaambati.svg" width="48" height="48" alt="@Kesavaambati" title="@Kesavaambati — 提出社区支持与容器镜像维护问题"></a>
</p>
每位 issue 或 PR 的作者都是 SILO 社区的一员,包括尚未合并的工作。已合并的修复、被采纳的方案和有效报告优先展示,其余参考首次参与时间;金色圆环突出经过审核的显著贡献。
[查看完整贡献记录](CONTRIBUTORS.md),了解每位贡献者的提案、修复与问题报告。
<a href="CONTRIBUTORS.md">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/contributors-dark.svg">
<img src="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/contributors-light.svg" alt="SILO 社区贡献者">
</picture>
</a>
[查看贡献记录与实际 PR 状态](CONTRIBUTORS.md)。
## Star History
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/star-history-dark.svg">
<img src="https://raw.githubusercontent.com/pgsty/silo/codex/repository-cards/star-history-light.svg" alt="SILO GitHub 星标历史">
</picture>
## 背景
+3 -3
View File
@@ -7,7 +7,7 @@ Silo-specific fixes or release notes.
## Supported Versions
Security fixes are tracked on the active development branch and summarized in
[docs/security/advisories.md](docs/security/advisories.md). Only the current
[the security advisory ledger](https://silo.pgsty.com/about/security-advisories/). Only the current
Silo release line is supported unless an advisory says otherwise.
## Inherited Fix Evidence
@@ -26,7 +26,7 @@ separately even when the fork preserves the original commit object and SHA.
The inherited [service-account](https://github.com/pgsty/silo/blob/c1a49490c78e9c3ebcad86ba0662319138ace190/cmd/admin-handlers-users_test.go#L211-L212)
and [STS](https://github.com/pgsty/silo/blob/c1a49490c78e9c3ebcad86ba0662319138ace190/cmd/sts-handlers_test.go#L45-L46)
regression groups remain part of `go test ./cmd`; see the
[canonical ledger](docs/security/advisories.md#inherited-upstream-advisory-baseline)
[canonical ledger](https://silo.pgsty.com/about/security-advisories/#inherited)
for the operator-facing record.
## Reporting a Vulnerability
@@ -42,4 +42,4 @@ For vulnerabilities in this fork:
## Disclosure Process
Fork-specific fixes and user-visible upgrade notes are published in [docs/security/advisories.md](docs/security/advisories.md). The fork-specific triage and remediation process is described in [VULNERABILITY_REPORT.md](VULNERABILITY_REPORT.md).
Fork-specific fixes and user-visible upgrade notes are published in [the security advisory ledger](https://silo.pgsty.com/about/security-advisories/). The fork-specific triage and remediation process is described in [VULNERABILITY_REPORT.md](VULNERABILITY_REPORT.md).
+1 -1
View File
@@ -34,4 +34,4 @@ Based on the report, the Silo maintainers investigate:
If the vulnerability exists in this fork itself, the maintainers will, when
feasible, fix the issue or implement reasonable countermeasures such that the
vulnerability can no longer be exploited. Fork-specific upgrade notes and
security advisories are published in `docs/security/advisories.md`.
security advisories are published in [the security advisory ledger](https://silo.pgsty.com/about/security-advisories/).
@@ -132,8 +132,8 @@
"MINIO_API_DELETE_CLEANUP_INTERVAL",
"MINIO_API_DISABLE_ODIRECT",
"MINIO_API_GZIP_OBJECTS",
"MINIO_API_LEGACY_BUCKET_RESOURCE_MATCH",
"MINIO_API_LIST_QUORUM",
"MINIO_API_MULTIPART_LISTING",
"MINIO_API_OBJECT_MAX_VERSIONS",
"MINIO_API_ODIRECT",
"MINIO_API_REMOTE_TRANSPORT_DEADLINE",
@@ -627,7 +627,6 @@
"x-minio-origin-endpoint",
"x-minio-prefixes-total",
"x-minio-read-quorum",
"x-minio-replication",
"x-minio-replication-actual-object-size",
"x-minio-replication-delete-status",
"x-minio-replication-deletemarker-status",
@@ -766,6 +765,7 @@
"/metrics/v3",
"/minio/grid/",
"/minio/grid/lock/",
"/multipart-preflight",
"/netperf",
"/notification",
"/oauth2/callback",
@@ -829,6 +829,7 @@
"/site-replication/peer/bucket-ops",
"/site-replication/peer/edit",
"/site-replication/peer/iam-item",
"/site-replication/peer/iam-revisions",
"/site-replication/peer/idp-settings",
"/site-replication/peer/join",
"/site-replication/peer/remove",
@@ -877,6 +878,7 @@
"/v2/metrics/cluster",
"/v2/metrics/node",
"/v2/metrics/resource",
"/v3/site-replication/peer/iam-revisions",
"/var/vcap/bosh",
"/verifybinary",
"/version",
@@ -932,6 +934,7 @@
"arn:minio:kms:::this-is-disregarded",
"arn:minio:kms:::xyz-test-key",
"arn:minio:replication:",
"arn:minio:replication::",
"arn:minio:replication::8320b6d18f9032b4700f1f03b50d8d1853de8f22cab86931ee794e12f190852c:destinationbucket",
"arn:minio:replication:::",
"arn:minio:replication:::dest-bucket",
+1 -2
View File
@@ -115,9 +115,8 @@ func collect(repo string) (manifest, error) {
fset := token.NewFileSet()
for _, rel := range files {
// Investigation artifacts contain synthetic routes and archived configurations.
// Migration notes and guard fixtures contain archived identifiers.
if rel == "SILO_REBRANDING_MIGRATION.md" ||
strings.HasPrefix(rel, "docs/investigations/") ||
strings.HasPrefix(rel, "buildscripts/rebrand-guard/") ||
strings.HasPrefix(rel, "buildscripts/helm-migration-guard/") {
continue
+1
View File
@@ -0,0 +1 @@
__pycache__/
+163
View File
@@ -0,0 +1,163 @@
# Copyright (c) 2026 Feng Ruohang
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
"""Pure, self-contained SVG rendering for SILO's README cards."""
from datetime import date, timedelta
from html import escape
import math
import xml.etree.ElementTree as ET
themes = {
'light': dict(bg='#ffffff', wash='#f2f7fc', edge='#d9e3ee', ink='#16222e',
muted='#62758a', blue='#1d588c', copper='#b4762e', grid='#e5edf5',
line='#2b6ca3', ring='#dce5ef', field='#f7f9fc', label='#3d4e61'),
'dark': dict(bg='#101923', wash='#152738', edge='#2b3c50', ink='#e8eef6',
muted='#93a3b8', blue='#7fb8e8', copper='#e0a35c', grid='#263749',
line='#5da2dd', ring='#3a4e63', field='#0b1119', label='#b6c2d2'),
}
def read_emblem(path):
emblem = ET.parse(path).getroot()
body = ''.join(ET.tostring(child, encoding='unicode') for child in emblem
if child.tag.rsplit('}', 1)[-1] in ('defs', 'g'))
return '\n'.join(line.rstrip() for line in body.splitlines()).strip()
def txt(x, y, value, size=14, color=None, weight=400, anchor='start', mono=False, spacing=None):
family = 'Menlo,Consolas,monospace' if mono else 'Arial,Helvetica,sans-serif'
extra = f' letter-spacing="{spacing}"' if spacing is not None else ''
return (f'<text x="{x}" y="{y}" font-family="{family}" font-size="{size}" '
f'font-weight="{weight}" fill="{color}" text-anchor="{anchor}"{extra}>'
f'{escape(str(value))}</text>')
def start(height, theme, title, description, emblem_body):
t = themes[theme]
return [f'<svg xmlns="http://www.w3.org/2000/svg" width="1000" height="{height}" '
f'viewBox="0 0 1000 {height}" role="img" aria-labelledby="title desc">',
f'<title id="title">{escape(title)}</title><desc id="desc">{escape(description)}</desc>',
'<defs><linearGradient id="surface" x1="0" y1="1" x2="1" y2="0">'
f'<stop offset="0" stop-color="{t["bg"]}"/>'
f'<stop offset="1" stop-color="{t["wash"]}"/></linearGradient>'
'<linearGradient id="accent" x1="0" y1="0" x2="1" y2="0">'
f'<stop offset="0" stop-color="{t["blue"]}"/>'
f'<stop offset="1" stop-color="{t["copper"]}"/></linearGradient>'
'<linearGradient id="area" x1="0" y1="0" x2="0" y2="1">'
f'<stop offset="0" stop-color="{t["line"]}" stop-opacity=".22"/>'
f'<stop offset="1" stop-color="{t["line"]}" stop-opacity=".015"/>'
'</linearGradient></defs>',
f'<rect x=".75" y=".75" width="998.5" height="{height-1.5}" rx="22" '
f'fill="url(#surface)" stroke="{t["edge"]}" stroke-width="1.5"/>',
f'<svg x="40" y="25" width="25" height="25" viewBox="230 213 570 570">{emblem_body}</svg>']
def heading(parts, t, eyebrow, title, subtitle, value, value_label):
parts.extend([
txt(76, 43, eyebrow, 11, t['muted'], 600, mono=True, spacing=1.7),
txt(40, 94, title, 32, t['ink'], 700),
txt(41, 123, subtitle, 14, t['muted']),
txt(958, 89, f'{value:,}', 45, t['ink'], 700, anchor='end'),
txt(957, 114, value_label, 10, t['muted'], 600, anchor='end', mono=True, spacing=1.5),
f'<path d="M40 146 H960" stroke="{t["edge"]}"/>',
])
def contributors(theme, people, snapshot, emblem_body):
height = 206 + 76 * math.ceil(len(people) / 10)
t = themes[theme]
parts = start(height, theme, f'SILO community — {len(people)} contributors',
f'The existing SILO community roll, including code, proposals and reports across related projects. '
f'Gold rings retain the existing significant-contribution designation. Snapshot {snapshot}.', emblem_body)
heading(parts, t, 'SILO / COMMUNITY', 'Contributors',
'Code, proposals & reports across SILO and related projects', len(people), 'COMMUNITY CONTRIBUTORS')
for row in range(math.ceil(len(people) / 10)):
group = people[row * 10:(row + 1) * 10]
row_width = len(group) * 91
for col, person in enumerate(group):
x = (1000 - row_width) / 2 + col * 91 + 45.5
y = 199 + row * 76
identifier = f'avatar-{row}-{col}'
featured = bool(person.get('featured'))
parts.append(f'<g><title>@{escape(person["handle"])} — {escape(person["what"])}</title>')
parts.append(f'<defs><clipPath id="{identifier}"><circle cx="{x}" cy="{y}" r="29"/></clipPath></defs>')
if featured:
parts.append(f'<circle cx="{x}" cy="{y}" r="34" fill="{t["copper"]}" opacity=".09"/>')
if person.get('avatarDataUrl'):
parts.append(f'<image x="{x-29}" y="{y-29}" width="58" height="58" '
f'clip-path="url(#{identifier})" href="{escape(person["avatarDataUrl"])}"/>')
else:
parts.append(f'<circle cx="{x}" cy="{y}" r="29" fill="{t["ring"]}"/>')
parts.append(txt(x, y + 9, person['handle'][0].upper(), 26, t['ink'], 700, 'middle'))
parts.append(f'<circle cx="{x}" cy="{y}" r="30.5" fill="none" '
f'stroke="{t["copper"] if featured else t["ring"]}" stroke-width="{2 if featured else 1.25}"/></g>')
parts.extend([
f'<path d="M40 {height-42} H960" stroke="{t["edge"]}"/>',
f'<circle cx="47" cy="{height-21}" r="4" fill="none" stroke="{t["copper"]}" stroke-width="1.5"/>',
txt(61, height-17, 'Gold rings mark significant contributions', 12, t['muted']),
txt(959, height-17, f'AS OF {snapshot}', 10, t['muted'], 500, 'end', mono=True, spacing=.6),
'</svg>',
])
return ''.join(parts)
def stars(theme, history, snapshot, emblem_body):
points = history['points']
star_count = points[-1]['stars']
t = themes[theme]
provenance = ('Initial history reconstructed · Daily totals since ' + history['bootstrap']['through'] + ' · UTC'
if history['bootstrap']['reconstructed'] else 'Observed daily star totals · UTC')
parts = start(558, theme, f'SILO star history — {star_count:,} stars',
f'GitHub repository pgsty/silo. {star_count:,} stars as of {snapshot}. ' +
provenance, emblem_body)
heading(parts, t, 'SILO / GITHUB', 'Star History', 'pgsty/silo', star_count, 'GITHUB STARS')
left, right, top, bottom = 76, 958, 177, 440
begin = date.fromisoformat(points[0]['date'])
end = date.fromisoformat(points[-1]['date'])
days = max(1, (end - begin).days)
maximum = max(500, math.ceil(max(p['stars'] for p in points) / 500) * 500)
tick_step = 10 ** max(0, int(math.log10(maximum)))
xy = lambda day, n: (left + (right-left)*(date.fromisoformat(day)-begin).days / days,
bottom-(bottom-top)*n/maximum)
for value in range(0, maximum+1, tick_step):
y = xy(points[0]['date'], value)[1]
parts.append(f'<path d="M{left} {y:.2f} H{right}" stroke="{t["grid"]}" stroke-dasharray="4 6"/>')
parts.append(txt(left-16, round(y+4, 2), f'{value / 1000:g}k' if value >= 1000 else str(value), 12, t['muted'], anchor='end'))
dates = sorted({begin + timedelta(days=round((end-begin).days*i/5)) for i in range(6)})
ticks = [(d.isoformat(), d.strftime('%b %Y') if days > 90 else d.strftime('%b %d')) for d in dates]
for day, label in ticks:
x = xy(day, 0)[0]
parts.append(f'<path d="M{x:.2f} {top} V{bottom}" stroke="{t["grid"]}" stroke-opacity=".65"/>')
anchor = 'start' if day == points[0]['date'] else 'end' if day == snapshot else 'middle'
parts.append(txt(round(x, 2), 466, label, 12, t['muted'], anchor=anchor))
coords = [xy(p['date'], p['stars']) for p in points]
line = 'M' + ' L'.join(f'{x:.2f} {y:.2f}' for x, y in coords)
area = line + f' L{coords[-1][0]:.2f} {bottom} L{left} {bottom} Z'
parts.extend([
f'<path d="{area}" fill="url(#area)"/>',
f'<path d="{line}" fill="none" stroke="url(#accent)" stroke-width="3" '
'stroke-linecap="round" stroke-linejoin="round"/>',
f'<path d="M{left} {bottom} H{right}" stroke="{t["edge"]}"/>',
])
x, y = coords[-1]
parts.extend([
f'<circle cx="{x}" cy="{y}" r="10" fill="{t["copper"]}" opacity=".12"/>',
f'<circle cx="{x}" cy="{y}" r="5" fill="{t["copper"]}" stroke="{t["bg"]}" stroke-width="2"/>',
f'<path d="M40 491 H960" stroke="{t["edge"]}"/>',
txt(40, 516, f'{begin:%b %Y} — {end:%b %Y}'.upper(), 10, t['muted'], 500, mono=True, spacing=.7),
txt(959, 516, f'SNAPSHOT {snapshot}', 10, t['muted'], 500, 'end', mono=True, spacing=.6),
txt(40, 539, provenance, 11, t['muted']),
'</svg>',
])
return ''.join(parts)
@@ -0,0 +1 @@
PyYAML==6.0.3
+178
View File
@@ -0,0 +1,178 @@
# Copyright (c) 2026 Feng Ruohang
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
"""Regression checks for historical accuracy, contributor scope and SVG safety."""
import base64
import json
from pathlib import Path
import tempfile
import unittest
from unittest.mock import patch
from urllib.error import URLError
import xml.etree.ElementTree as ET
import render
import update
NS = {'s': 'http://www.w3.org/2000/svg'}
PNG = base64.b64decode('iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mP8/x8AAwMCAO+a7mgAAAAASUVORK5CYII=')
def person(handle='Alice', group='reports', featured=False):
return {'handle': handle, 'group': group, 'featured': featured,
'what': 'A reviewed contribution', 'firstContribution': '2026-09-01'}
def history():
return {'repository': 'pgsty/silo',
'bootstrap': {'through': '2026-09-15', 'reconstructed': True},
'points': [{'date': '2026-09-14', 'stars': 100}, {'date': '2026-09-15', 'stars': 105}]}
class HistoryTests(unittest.TestCase):
def test_new_day_preserves_old_counts_and_unstars(self):
before = history()
after = update.update_history(before, '2026-09-16', 103)
self.assertEqual(after['points'][:-1], before['points'])
self.assertEqual(after['points'][-1], {'date': '2026-09-16', 'stars': 103})
self.assertEqual(before, history())
def test_same_day_rerun_replaces_instead_of_appending(self):
first = update.update_history(history(), '2026-09-15', 107)
self.assertEqual(len(first['points']), 2)
self.assertEqual(first, update.update_history(first, '2026-09-15', 107))
def test_missing_days_are_not_invented(self):
result = update.update_history(history(), '2026-09-18', 106)
self.assertEqual([p['date'] for p in result['points']], ['2026-09-14', '2026-09-15', '2026-09-18'])
def test_rejects_wrong_repository_and_corrupt_history(self):
cases = []
wrong = history(); wrong['repository'] = 'someone/else'; cases.append(wrong)
duplicate = history(); duplicate['points'].append(duplicate['points'][-1]); cases.append(duplicate)
unordered = history(); unordered['points'].reverse(); cases.append(unordered)
negative = history(); negative['points'][0]['stars'] = -1; cases.append(negative)
future = history(); future['points'][-1]['date'] = '2026-09-20'; cases.append(future)
for case in cases:
with self.subTest(case=case), self.assertRaises(ValueError):
update.update_history(case, '2026-09-16', 100)
def test_first_run_has_no_fabricated_history(self):
result = update.update_history(None, '2026-09-16', 10)
self.assertFalse(result['bootstrap']['reconstructed'])
self.assertEqual(result['points'], [{'date': '2026-09-16', 'stars': 10}])
class ContributorTests(unittest.TestCase):
def test_retries_truncated_json_before_using_it(self):
with patch('update.request', side_effect=[b'{"partial":', b'{"ok":true}']), patch('update.time.sleep'):
self.assertEqual(update.GitHub('').get('repos/pgsty/silo'), {'ok': True})
def test_paginates_past_one_full_page(self):
class API(update.GitHub):
def __init__(self): self.calls = []
def get(self, path):
self.calls.append(path)
return list(range(100)) if 'page=1&' in path else [100]
api = API()
self.assertEqual(len(list(api.issues('pgsty/silo'))), 101)
self.assertIn('state=all', api.calls[0])
self.assertIn('page=2&', api.calls[1])
def test_bots_deduplication_unmerged_work_and_reviewed_credit(self):
def issue(login, kind='issue', user_type='User'):
item = {'user': {'login': login, 'type': user_type, 'avatar_url': ''}, 'created_at': '2026-09-02T00:00:00Z'}
if kind != 'issue': item['pull_request'] = {'merged_at': None if kind == 'open' else '2026-09-03T00:00:00Z'}
return item
class API:
def issues(self, _repo):
return [issue('alice'), issue('Bob', 'open'), issue('Bob', 'merged'),
issue('Carol', 'open'), issue('Copilot'), issue('robot', user_type='Bot')]
curated = {'repositories': ['pgsty/silo', 'pgsty/mc'], 'bots': ['Copilot'],
'people': [person('Alice', featured=True), person('Reporter'), person('Copilot')]}
result = update.collect_people(API(), curated)
self.assertEqual({p['handle'] for p in result}, {'Alice', 'Bob', 'Carol', 'Reporter'})
self.assertEqual(result[0]['handle'], 'Bob')
self.assertEqual(result[1]['handle'], 'Carol')
self.assertTrue(next(p for p in result if p['handle'] == 'Alice')['featured'])
self.assertEqual(next(p for p in result if p['handle'] == 'Bob')['group'], 'code')
self.assertFalse(next(p for p in result if p['handle'] == 'Carol')['featured'])
def test_newer_reviewed_preview_survives_until_site_catches_up(self):
remote = {'updated': '2026-09-16T03:00:00+00:00'}
cached = {'updated': '2026-09-16T04:00:00+00:00'}
self.assertIs(update.select_curated(remote, cached), cached)
newer = {'updated': '2026-09-17T03:00:00+00:00'}
self.assertIs(update.select_curated(newer, cached), newer)
def test_avatar_failure_reuses_raster_cache(self):
previous = {**person(), 'avatarDataUrl': update.raster_data_url(PNG)}
with patch('update.request', side_effect=URLError('unavailable')):
result = update.add_avatars(None, [{**person(), 'avatarUrl': 'https://avatars.githubusercontent.com/u/1'}], [previous])
self.assertEqual(result[0]['avatarDataUrl'], previous['avatarDataUrl'])
with self.assertRaises(ValueError): update.raster_data_url(b'<svg onload="bad()"/>')
with self.assertRaises(ValueError): update.cached_avatar({'avatarDataUrl': 'data:image/svg+xml;base64,PHN2Zy8+'})
def test_fetch_failure_leaves_published_assets_untouched(self):
class API:
def get(self, path):
if path == 'repos/pgsty/silo': return {'full_name': 'pgsty/silo', 'stargazers_count': 106}
raise URLError('roster unavailable')
with tempfile.TemporaryDirectory() as directory:
out = Path(directory)
original = json.dumps(history())
(out / 'history.json').write_text(original)
(out / 'contributors-light.svg').write_text('previous image')
with self.assertRaises(URLError): update.refresh(out, API(), Path('.'))
self.assertEqual((out / 'history.json').read_text(), original)
self.assertEqual((out / 'contributors-light.svg').read_text(), 'previous image')
class RenderTests(unittest.TestCase):
def test_real_emblem_generates_clean_xml(self):
emblem = render.read_emblem(Path(__file__).resolve().parents[2] / '.github/silo.svg')
svg = render.contributors('light', [person()], '2026-09-16', emblem)
ET.fromstring(svg)
self.assertTrue(all(line == line.rstrip() for line in svg.splitlines()))
def test_all_avatars_fit_when_the_roster_grows(self):
people = [{**person(f'person-{i}'), 'avatarDataUrl': update.raster_data_url(PNG)} for i in range(151)]
for theme in ('light', 'dark'):
root = ET.fromstring(render.contributors(theme, people, '2026-09-16', ''))
images = root.findall('.//s:image', NS)
self.assertEqual(len(images), 151)
footer = float(root.attrib['height']) - 42
self.assertTrue(all(float(i.attrib['y']) + float(i.attrib['height']) < footer for i in images))
self.assertTrue(all(i.attrib['href'].startswith('data:image/png;base64,') for i in images))
def test_untrusted_text_is_escaped(self):
data = [{**person(), 'what': '<script>alert("x")</script> & contributions'}]
svg = render.contributors('light', data, '2026-09-16', '')
root = ET.fromstring(svg)
self.assertEqual(root.findall('.//s:script', NS), [])
self.assertIn('&lt;script&gt;', svg)
def test_single_point_and_decreasing_star_history_render(self):
for data in (update.update_history(None, '2026-09-16', 0), update.update_history(history(), '2026-09-16', 90)):
for theme in ('light', 'dark'):
svg = render.stars(theme, data, '2026-09-16', '')
ET.fromstring(svg)
self.assertNotIn('nan', svg.lower())
self.assertNotIn('inf', svg.lower())
if __name__ == '__main__':
unittest.main()
+295
View File
@@ -0,0 +1,295 @@
#!/usr/bin/env python3
# Copyright (c) 2026 Feng Ruohang
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
"""Refresh the generated-asset checkout; publishing is handled by the workflow."""
import argparse
import base64
from concurrent.futures import ThreadPoolExecutor
from datetime import date, datetime, timezone
import json
from http.client import IncompleteRead
import os
from pathlib import Path
import re
import sys
import time
from urllib.error import HTTPError, URLError
from urllib.parse import urlparse
from urllib.request import Request, urlopen
import xml.etree.ElementTree as ET
import yaml
import render
REPOSITORY = 'pgsty/silo'
SOURCE = 'repos/pgsty/silo.pgsty.com/contents/data/home/contributors.yaml?ref=main'
GROUPS = ('code', 'proposed', 'reports')
HANDLE = re.compile(r'[A-Za-z0-9][A-Za-z0-9-]{0,38}\Z')
def request(url, token='', limit=8 * 1024 * 1024):
headers = {'User-Agent': 'silo-repository-cards', 'Accept': 'application/vnd.github+json'}
if urlparse(url).netloc == 'api.github.com':
headers['X-GitHub-Api-Version'] = '2022-11-28'
if token:
headers['Authorization'] = f'Bearer {token}'
for attempt in range(3):
try:
with urlopen(Request(url, headers=headers), timeout=25) as response:
data = response.read(limit + 1)
if len(data) > limit:
raise ValueError('Response exceeds the size limit')
expected = response.headers.get('Content-Length')
if expected is not None and len(data) != int(expected):
raise URLError('Incomplete response body')
return data
except HTTPError as exc:
if exc.code < 500 or attempt == 2:
raise
except (URLError, TimeoutError, IncompleteRead):
if attempt == 2:
raise
time.sleep(attempt + 1)
class GitHub:
def __init__(self, token):
self.token = token
def get(self, path):
for attempt in range(3):
try:
return json.loads(request('https://api.github.com/' + path, self.token))
except (json.JSONDecodeError, UnicodeDecodeError) as exc:
if attempt == 2:
raise ValueError(f'Incomplete or invalid GitHub JSON: {path}') from exc
time.sleep(attempt + 1)
def issues(self, repository):
page = 1
while True:
batch = self.get(f'repos/{repository}/issues?state=all&per_page=100&page={page}&sort=created&direction=asc')
if not isinstance(batch, list):
raise ValueError(f'Invalid issues response for {repository}')
yield from batch
if len(batch) < 100:
return
page += 1
def curated_snapshot(data, revision):
updated = str(data['updated'])
datetime.fromisoformat(updated)
repositories = [item['repo'] for item in data['repositories']]
if not repositories or any(not re.fullmatch(r'pgsty/[A-Za-z0-9_.-]+', repo) for repo in repositories):
raise ValueError('Invalid contributor repository scope')
people = []
for group in GROUPS:
for entry in data[group]:
if not HANDLE.fullmatch(entry['handle']):
raise ValueError('Invalid GitHub contributor handle')
people.append({
'handle': entry['handle'], 'group': group,
'featured': bool(entry.get('featured')), 'what': entry['what'],
'firstContribution': str(entry.get('firstContribution', '9999-12-31')),
})
if not people or len({p['handle'].lower() for p in people}) != len(people):
raise ValueError('Empty or duplicate contributor roster')
return {'updated': updated, 'revision': revision, 'repositories': repositories,
'bots': data.get('bots', ['Copilot', 'dependabot[bot]']), 'people': people}
def select_curated(remote, cached):
# The initial, approved preview can contain reviewed credit not published by
# the companion site yet. Keep that newer snapshot until the site catches up.
if cached and datetime.fromisoformat(cached['updated']) > datetime.fromisoformat(remote['updated']):
return cached
return remote
def collect_people(api, curated):
bots = {name.lower() for name in curated['bots']}
people = {p['handle'].lower(): dict(p) for p in curated['people']
if p['handle'].lower() not in bots and not p['handle'].lower().endswith('[bot]')}
order = {p['handle'].lower(): index for index, p in enumerate(curated['people'])}
for repository in curated['repositories']:
print(f'Reading issue and PR authors: {repository}', flush=True)
for issue in api.issues(repository):
user = issue.get('user') or {}
handle = user.get('login', '')
key = handle.lower()
if user.get('type') != 'User' or key in bots or key.endswith('[bot]'):
continue
if not HANDLE.fullmatch(handle):
raise ValueError('Invalid issue author')
pr = issue.get('pull_request')
group = 'code' if pr and pr.get('merged_at') else 'proposed' if pr else 'reports'
first = issue['created_at'][:10]
date.fromisoformat(first)
person = people.setdefault(key, {
'handle': handle, 'group': group, 'featured': False,
'what': 'Contributed an issue or pull request to SILO and related projects',
'firstContribution': first,
})
person['avatarUrl'] = user.get('avatar_url', '')
person['firstContribution'] = min(person['firstContribution'], first)
if GROUPS.index(group) < GROUPS.index(person['group']):
person['group'] = group
if not people:
raise ValueError('No human contributors were collected')
return sorted(people.values(), key=lambda p: (
GROUPS.index(p['group']), not p['featured'],
order.get(p['handle'].lower(), len(order)), p['firstContribution'], p['handle'].lower()))
def raster_data_url(data):
if data.startswith(b'\x89PNG\r\n\x1a\n'):
mime = 'image/png'
elif data.startswith(b'\xff\xd8\xff'):
mime = 'image/jpeg'
elif data.startswith((b'GIF87a', b'GIF89a')):
mime = 'image/gif'
elif data[:4] == b'RIFF' and data[8:12] == b'WEBP':
mime = 'image/webp'
else:
raise ValueError('Avatar is not a raster image')
return f'data:{mime};base64,' + base64.b64encode(data).decode('ascii')
def cached_avatar(person):
value = person.get('avatarDataUrl', '')
if not value:
return ''
prefix, encoded = value.split(',', 1)
if prefix not in ('data:image/png;base64', 'data:image/jpeg;base64', 'data:image/gif;base64', 'data:image/webp;base64'):
raise ValueError('Invalid cached avatar format')
raw = base64.b64decode(encoded, validate=True)
if len(raw) > 512 * 1024 or raster_data_url(raw) != value:
raise ValueError('Invalid cached avatar')
return value
def add_avatars(api, people, previous):
cached = {p['handle'].lower(): cached_avatar(p) for p in previous}
def update(person):
person = dict(person)
try:
url = person.pop('avatarUrl', '') or api.get('users/' + person['handle'])['avatar_url']
parsed = urlparse(url)
if parsed.scheme != 'https' or parsed.netloc != 'avatars.githubusercontent.com':
raise ValueError('Unexpected avatar host')
data = request(url + ('&' if '?' in url else '?') + 's=96', limit=512 * 1024)
person['avatarDataUrl'] = raster_data_url(data)
except (HTTPError, URLError, TimeoutError, IncompleteRead, ValueError, KeyError) as exc:
person.pop('avatarUrl', None)
person['avatarDataUrl'] = cached.get(person['handle'].lower(), '')
print(f'Avatar fallback for @{person["handle"]}: {type(exc).__name__}', file=sys.stderr)
return person
with ThreadPoolExecutor(max_workers=6) as pool:
return list(pool.map(update, people))
def update_history(history, day, stars):
date.fromisoformat(day)
if type(stars) is not int or stars < 0:
raise ValueError('Invalid repository star count')
if history is None:
history = {'repository': REPOSITORY, 'bootstrap': {'through': day, 'reconstructed': False}, 'points': []}
if history['repository'] != REPOSITORY:
raise ValueError('Star history belongs to a different repository')
date.fromisoformat(history['bootstrap']['through'])
dates = []
for point in history['points']:
date.fromisoformat(point['date'])
if type(point['stars']) is not int or point['stars'] < 0:
raise ValueError('Invalid historical star count')
dates.append(point['date'])
if dates != sorted(set(dates)) or any(d > day for d in dates):
raise ValueError('History contains duplicate, unordered, or future dates')
# Replace today's observation, preserve previous days, and allow unstars.
points = [dict(p) for p in history['points'] if p['date'] != day]
points.append({'date': day, 'stars': stars})
return {**history, 'points': points}
def read_json(path, default=None):
return json.loads(path.read_text()) if path.exists() else default
def refresh(output, api, source_root):
day = datetime.now(timezone.utc).date().isoformat()
metadata = api.get('repos/' + REPOSITORY)
if metadata['full_name'].lower() != REPOSITORY:
raise ValueError('Unexpected repository metadata')
history = update_history(read_json(output / 'history.json'), day, metadata['stargazers_count'])
source = api.get(SOURCE)
reviewed = yaml.safe_load(base64.b64decode(source['content'], validate=False))
curated = select_curated(curated_snapshot(reviewed, source['sha']), read_json(output / 'curated.json'))
people = collect_people(api, curated)
previous = read_json(output / 'contributors.json', {}).get('people', [])
people = add_avatars(api, people, previous)
emblem = render.read_emblem(source_root / '.github/silo.svg')
payloads = {}
for theme in ('light', 'dark'):
payloads[f'contributors-{theme}.svg'] = render.contributors(theme, people, day, emblem) + '\n'
payloads[f'star-history-{theme}.svg'] = render.stars(theme, history, day, emblem) + '\n'
for svg in payloads.values():
ET.fromstring(svg)
for name, data in {
'history.json': history,
'curated.json': curated,
'contributors.json': {'repository': REPOSITORY, 'updated': day, 'people': people},
}.items():
payloads[name] = json.dumps(data, indent=2, ensure_ascii=False) + '\n'
payloads['README.md'] = f'''# SILO repository cards
Generated by [Repository Cards](https://github.com/pgsty/silo/actions/workflows/repository-cards.yml)
at 00:00 UTC daily (08:00 Asia/Shanghai). GitHub may queue scheduled runs.
Snapshot: {day}. {metadata['stargazers_count']:,} stars; {len(people)} community contributors.
- `contributors-light.svg` / `contributors-dark.svg`: human issue and PR authors across the SILO project scope, plus reviewed acknowledgements. Bots are excluded. Gold rings follow the reviewed companion-site roster; new authors are collected automatically.
- `star-history-light.svg` / `star-history-dark.svg`: initial history reconstructed from the then-current stargazers; later points are daily observed totals, including decreases. Missing days are not fabricated.
- `curated.json`: a cache of reviewed contributor credit from `pgsty/silo.pgsty.com/data/home/contributors.yaml`. The approved initial preview may be newer than the published site; a newer reviewed snapshot is retained until the site catches up.
- `contributors.json`: generated contributor data and embedded raster avatars. Failed avatar refreshes use the previous image, or an initial when no image is available.
- `history.json`: persistent daily totals. Keep this file when regenerating images.
The SVGs are self-contained. Source and instructions live on the default branch;
this branch contains generated assets only. Do not merge it into `main`.
'''
# Collect and validate everything before touching the publication checkout.
output.mkdir(parents=True, exist_ok=True)
for filename, text in payloads.items():
(output / filename).write_text(text)
print(f'{day}: {len(people)} contributors; {metadata["stargazers_count"]:,} stars; {len(history["points"])} history points')
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--output', type=Path, required=True)
args = parser.parse_args()
configured = os.environ.get('GITHUB_REPOSITORY', REPOSITORY)
if configured.lower() != REPOSITORY:
raise SystemExit('This workflow is scoped to pgsty/silo')
refresh(args.output, GitHub(os.environ.get('GH_TOKEN', '')), Path(__file__).resolve().parents[2])
if __name__ == '__main__':
main()
+2
View File
@@ -163,6 +163,7 @@ func registerAdminRouter(router *mux.Router, enableConfigOps bool) {
// StorageInfo operations
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/storageinfo").HandlerFunc(adminMiddleware(adminAPI.StorageInfoHandler, traceAllFlag))
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/multipart-preflight").HandlerFunc(adminMiddleware(adminAPI.MultipartPreflightHandler, traceAllFlag))
// DataUsageInfo operations
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/datausageinfo").HandlerFunc(adminMiddleware(adminAPI.DataUsageInfoHandler, traceAllFlag))
// Metrics operation
@@ -388,6 +389,7 @@ func registerAdminRouter(router *mux.Router, enableConfigOps bool) {
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/join").HandlerFunc(adminMiddleware(adminAPI.SRPeerJoin))
adminRouter.Methods(http.MethodPut).Path(adminVersion+"/site-replication/peer/bucket-ops").HandlerFunc(adminMiddleware(adminAPI.SRPeerBucketOps)).Queries("bucket", "{bucket:.*}").Queries("operation", "{operation:.*}")
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-item").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateIAMItem))
adminRouter.Methods(http.MethodGet, http.MethodPut).Path(adminVersion + "/site-replication/peer/iam-revisions").HandlerFunc(adminMiddleware(adminAPI.SRPeerIAMRevisions))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/peer/bucket-meta").HandlerFunc(adminMiddleware(adminAPI.SRPeerReplicateBucketItem))
adminRouter.Methods(http.MethodGet).Path(adminVersion + "/site-replication/peer/idp-settings").HandlerFunc(adminMiddleware(adminAPI.SRPeerGetIDPSettings))
adminRouter.Methods(http.MethodPut).Path(adminVersion + "/site-replication/edit").HandlerFunc(adminMiddleware(adminAPI.SiteReplicationEdit))
+24
View File
@@ -450,6 +450,9 @@ const (
ErrAdminNoSecretKey
ErrIAMNotInitialized
ErrMultipartListingLegacy
ErrMultipartListingIdentity
ErrSlowDown
apiErrCodeEnd // This is used only for the testing code
)
@@ -1336,6 +1339,21 @@ var errorCodes = errorCodeMap{
Description: "IAM sub-system not initialized yet, please try again.",
HTTPStatusCode: http.StatusServiceUnavailable,
},
ErrMultipartListingLegacy: {
Code: "MultipartListingNotReady",
Description: "Legacy multipart uploads prevent a complete listing. Upgrade all writers, drain old uploads and run the multipart preflight check.",
HTTPStatusCode: http.StatusServiceUnavailable,
},
ErrMultipartListingIdentity: {
Code: "MultipartListingMetadataInvalid",
Description: "Multipart upload metadata is inconsistent. Run the multipart preflight check to locate the affected storage set.",
HTTPStatusCode: http.StatusServiceUnavailable,
},
ErrSlowDown: {
Code: "SlowDown",
Description: "Please reduce your request rate",
HTTPStatusCode: http.StatusServiceUnavailable,
},
ErrBucketMetadataNotInitialized: {
Code: "XMinioBucketMetadataNotInitialized",
Description: "Bucket metadata not initialized yet, please try again.",
@@ -2173,6 +2191,10 @@ func toAPIErrorCode(ctx context.Context, err error) (apiErr APIErrorCode) {
err = unwrapAll(err)
switch err {
case errMultipartListingLegacy:
apiErr = ErrMultipartListingLegacy
case errMultipartListingIdentity:
apiErr = ErrMultipartListingIdentity
case errCompleteMultipartChecksumMismatch, errCompleteMultipartChecksumTypeMismatch:
apiErr = ErrBadDigest
case errMissingPartChecksum:
@@ -2303,6 +2325,8 @@ func toAPIErrorCode(ctx context.Context, err error) (apiErr APIErrorCode) {
}
switch err.(type) {
case SlowDown:
apiErr = ErrSlowDown
case StorageFull:
apiErr = ErrStorageFull
case hash.BadDigest:
+1 -1
View File
@@ -41,7 +41,7 @@ import (
const (
maxObjectList = 1000 // Limit number of objects in a listObjectsResponse/listObjectsVersionsResponse.
maxDeleteList = 1000 // Limit number of objects deleted in a delete call.
maxUploadsList = 10000 // Limit number of uploads in a listUploadsResponse.
maxUploadsList = 1000 // Limit number of uploads in a listUploadsResponse.
maxPartsList = 10000 // Limit number of parts in a listPartsResponse.
)
File diff suppressed because one or more lines are too long
-8
View File
@@ -277,14 +277,6 @@ func (api objectAPIHandlers) ListMultipartUploadsHandler(w http.ResponseWriter,
return
}
if keyMarker != "" {
// Marker not common with prefix is not implemented.
if !HasPrefix(keyMarker, prefix) {
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(ErrNotImplemented), r.URL)
return
}
}
listMultipartsInfo, err := objectAPI.ListMultipartUploads(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
if err != nil {
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
+3 -3
View File
@@ -425,7 +425,7 @@ func testListMultipartUploadsHandler(obj ObjectLayer, instanceType, bucketName s
shouldPass: true,
},
// Test case - 4.
// Setting Invalid prefix and marker combination.
// A key marker outside the prefix is valid and produces an empty page.
{
bucket: bucketName,
prefix: "asia",
@@ -435,8 +435,8 @@ func testListMultipartUploadsHandler(obj ObjectLayer, instanceType, bucketName s
maxUploads: "0",
accessKey: credentials.AccessKey,
secretKey: credentials.SecretKey,
expectedRespStatus: http.StatusNotImplemented,
shouldPass: false,
expectedRespStatus: http.StatusOK,
shouldPass: true,
},
// Test case - 5.
// Invalid upload id and marker combination.
+3
View File
@@ -418,6 +418,9 @@ func getReplicationState(rinfos replicatedInfos, prevState ReplicationState, vID
for _, rinfo := range rinfos.Targets {
if rinfo.ResyncTimestamp != "" {
if rs.ResetStatusesMap == nil {
rs.ResetStatusesMap = make(map[string]string)
}
rs.ResetStatusesMap[targetResetHeader(rinfo.Arn)] = rinfo.ResyncTimestamp
}
}
+111 -63
View File
@@ -425,6 +425,7 @@ func checkReplicateDelete(ctx context.Context, bucket string, dobj ObjectToDelet
// the mere presence or absence of the target version.
func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, objectAPI ObjectLayer) replicatedInfos {
var replicationStatus replication.StatusType
isPurge := dobj.isVersionPurge()
bucket := dobj.Bucket
versionID := dobj.DeleteMarkerVersionID
if versionID == "" {
@@ -484,6 +485,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
lk := objectAPI.NewNSLock(bucket, "/[replicate]/"+dobj.ObjectName)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
dobj.RetryCount++
globalReplicationPool.Get().queueMRFSave(dobj.ToMRFEntry())
sendEvent(eventArgs{
BucketName: bucket,
@@ -546,25 +548,38 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
replicationStatus = rinfos.ReplicationStatus()
prevStatus := dobj.DeleteMarkerReplicationStatus()
if dobj.VersionID != "" {
prevStatus = replication.StatusType(dobj.VersionPurgeStatus())
replicationStatus = replication.StatusType(rinfos.VersionPurgeStatus())
if isPurge {
prevStatus = purgeReplicationStatus(dobj.VersionPurgeStatus())
replicationStatus = purgeReplicationStatus(rinfos.VersionPurgeStatus())
}
// to decrement pending count later.
for _, rinfo := range rinfos.Targets {
if rinfo.ReplicationStatus != rinfo.PrevReplicationStatus {
globalReplicationStats.Load().Update(dobj.Bucket, rinfo, replicationStatus,
prevStatus)
status, previous := rinfo.ReplicationStatus, rinfo.PrevReplicationStatus
if isPurge {
status = purgeReplicationStatus(rinfo.VersionPurgeStatus)
previous = purgeReplicationStatus(dobj.ReplicationState.PurgeTargets[rinfo.Arn])
}
if status != previous {
globalReplicationStats.Load().Update(dobj.Bucket, rinfo, status, previous)
}
}
eventName := event.ObjectReplicationComplete
if replicationStatus == replication.Failed {
eventName = event.ObjectReplicationFailed
dobj.RetryCount++
globalReplicationPool.Get().queueMRFSave(dobj.ToMRFEntry())
}
drs := getReplicationState(rinfos, dobj.ReplicationState, dobj.VersionID)
if isPurge {
// A purge must not rewrite the marker's creation/replica metadata.
// Multiple serialized empty target statuses can parse as nonempty,
// so explicitly send an empty creation update to the metadata writer.
drs.ReplicationStatusInternal = ""
drs.Targets = nil
drs.ReplicaStatus = ""
}
if replicationStatus != prevStatus {
drs.ReplicationTimeStamp = UTCNow()
}
@@ -606,6 +621,7 @@ func replicateDelete(ctx context.Context, dobj DeletedObjectReplicationInfo, obj
}
func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationInfo, tgt *TargetClient) (rinfo replicatedTargetInfo) {
isPurge := dobj.isVersionPurge()
versionID := dobj.DeleteMarkerVersionID
if versionID == "" {
versionID = dobj.VersionID
@@ -615,42 +631,50 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
rinfo.OpType = dobj.OpType
rinfo.endpoint = tgt.EndpointURL().Host
rinfo.secure = tgt.EndpointURL().Scheme == "https"
// A purge leaves ReplicationStatus empty: the metadata writer interprets
// that as preserving the entire creation-status block, including targets
// outside this fan-out. Only VersionPurgeStatus records the purge outcome.
defer func() {
if rinfo.ReplicationStatus == replication.Completed && tgt.ResetID != "" && dobj.OpType == replication.ExistingObjectReplicationType {
completed := rinfo.ReplicationStatus == replication.Completed
if isPurge {
completed = rinfo.VersionPurgeStatus == replication.VersionPurgeComplete
}
if completed && tgt.ResetID != "" && dobj.OpType == replication.ExistingObjectReplicationType {
rinfo.ResyncTimestamp = fmt.Sprintf("%s;%s", UTCNow().Format(http.TimeFormat), tgt.ResetID)
}
}()
if dobj.VersionID == "" && rinfo.PrevReplicationStatus == replication.Completed && dobj.OpType != replication.ExistingObjectReplicationType {
if !isPurge && rinfo.PrevReplicationStatus == replication.Completed && dobj.OpType != replication.ExistingObjectReplicationType {
rinfo.ReplicationStatus = rinfo.PrevReplicationStatus
return rinfo
}
if dobj.VersionID != "" && rinfo.VersionPurgeStatus == replication.VersionPurgeComplete {
if isPurge && rinfo.VersionPurgeStatus == replication.VersionPurgeComplete {
return rinfo
}
if globalBucketTargetSys.isOffline(tgt.EndpointURL()) {
replLogOnceIf(ctx, fmt.Errorf("remote target is offline for bucket:%s arn:%s", dobj.Bucket, tgt.ARN), "replication-target-offline-delete-"+tgt.ARN)
rinfo.Err = fmt.Errorf("remote target is offline for bucket:%s arn:%s", dobj.Bucket, tgt.ARN)
replLogOnceIf(ctx, rinfo.Err, "replication-target-offline-delete-"+tgt.ARN)
sendEvent(eventArgs{
BucketName: dobj.Bucket,
Object: ObjectInfo{
Bucket: dobj.Bucket,
Name: dobj.ObjectName,
VersionID: dobj.VersionID,
VersionID: versionID,
DeleteMarker: dobj.DeleteMarker,
},
UserAgent: "Internal: [Replication]",
Host: globalLocalNodeName,
EventName: event.ObjectReplicationNotTracked,
})
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Failed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeFailed
} else {
rinfo.ReplicationStatus = replication.Failed
}
return rinfo
}
// early return if already replicated delete marker for existing object replication/ healing delete markers
if dobj.DeleteMarkerVersionID != "" {
if !isPurge && dobj.DeleteMarkerVersionID != "" {
toi, err := tgt.StatObject(ctx, tgt.Bucket, dobj.ObjectName, minio.StatObjectOptions{
VersionID: versionID,
Internal: minio.AdvancedGetOptions{
@@ -662,16 +686,10 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
switch {
case isErrMethodNotAllowed(serr):
// delete marker already replicated
if dobj.VersionID == "" && rinfo.VersionPurgeStatus.Empty() {
rinfo.ReplicationStatus = replication.Completed
return rinfo
}
rinfo.ReplicationStatus = replication.Completed
return rinfo
case isErrObjectNotFound(serr), isErrVersionNotFound(serr):
// version being purged is already not found on target.
if !rinfo.VersionPurgeStatus.Empty() {
rinfo.VersionPurgeStatus = replication.VersionPurgeComplete
return rinfo
}
// The marker still needs to be created on the target.
case isErrReadQuorum(serr), isErrWriteQuorum(serr):
// destination has some quorum issues, perform removeObject() anyways
// to complete the operation.
@@ -691,7 +709,7 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
rmErr := tgt.RemoveObject(ctx, tgt.Bucket, dobj.ObjectName, minio.RemoveObjectOptions{
VersionID: versionID,
Internal: minio.AdvancedRemoveOptions{
ReplicationDeleteMarker: dobj.DeleteMarkerVersionID != "",
ReplicationDeleteMarker: !isPurge && dobj.DeleteMarkerVersionID != "",
ReplicationMTime: dobj.DeleteMarkerMTime.Time,
ReplicationStatus: minio.ReplicationStatusReplica,
ReplicationRequest: true, // always set this to distinguish between `mc mirror` replication and serverside
@@ -699,20 +717,20 @@ func replicateDeleteToTarget(ctx context.Context, dobj DeletedObjectReplicationI
})
if rmErr != nil {
rinfo.Err = rmErr
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Failed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeFailed
} else {
rinfo.ReplicationStatus = replication.Failed
}
replLogIf(ctx, fmt.Errorf("unable to replicate delete marker to %s: %s/%s(%s): %w", tgt.EndpointURL(), tgt.Bucket, dobj.ObjectName, versionID, rmErr))
if rmErr != nil && minio.IsNetworkOrHostDown(rmErr, true) && !globalBucketTargetSys.isOffline(tgt.EndpointURL()) {
globalBucketTargetSys.markOffline(tgt.EndpointURL())
}
} else {
if dobj.VersionID == "" {
rinfo.ReplicationStatus = replication.Completed
} else {
if isPurge {
rinfo.VersionPurgeStatus = replication.VersionPurgeComplete
} else {
rinfo.ReplicationStatus = replication.Completed
}
}
return rinfo
@@ -781,6 +799,18 @@ func (m caseInsensitiveMap) Lookup(key string) (string, bool) {
return "", false
}
// replicationTaggingTimestamp carries a recorded removal even when tags are
// empty. Only legacy nonempty tags use ModTime; absence is not a tombstone.
func replicationTaggingTimestamp(objInfo ObjectInfo) (time.Time, error) {
if stamp, ok := caseInsensitiveMap(objInfo.UserDefined).Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
return time.Parse(time.RFC3339Nano, stamp)
}
if objInfo.UserTags != "" {
return objInfo.ModTime, nil
}
return time.Time{}, nil
}
func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (putOpts minio.PutObjectOptions, isMP bool, err error) {
meta := make(map[string]string)
isSSEC := crypto.SSEC.IsEncrypted(objInfo.UserDefined)
@@ -850,17 +880,12 @@ func putReplicationOpts(ctx context.Context, sc string, objInfo ObjectInfo) (put
tag, _ := tags.ParseObjectTags(objInfo.UserTags)
if tag != nil {
putOpts.UserTags = tag.ToMap()
// set tag timestamp in opts
tagTimestamp := objInfo.ModTime
if tagTmstampStr, ok := objInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]; ok {
tagTimestamp, err = time.Parse(time.RFC3339Nano, tagTmstampStr)
if err != nil {
return putOpts, false, err
}
}
putOpts.Internal.TaggingTimestamp = tagTimestamp
}
}
putOpts.Internal.TaggingTimestamp, err = replicationTaggingTimestamp(objInfo)
if err != nil {
return putOpts, false, err
}
lkMap := caseInsensitiveMap(objInfo.UserDefined)
if lang, ok := lkMap.Lookup(xhttp.ContentLanguage); ok {
@@ -1003,6 +1028,12 @@ func getReplicationAction(oi1 ObjectInfo, oi2 minio.ObjectInfo, opType replicati
if (oi2.UserTagCount > 0 && !reflect.DeepEqual(oi2Map, t.ToMap())) || (oi2.UserTagCount != len(t.ToMap())) {
return replicateMetadata
}
// HEAD does not report the tag revision. Equal values can hide a newer
// deletion or re-addition, so scheduled metadata/heal work must deliver it.
// Completed objects are still excluded by the existing scanner gates.
if _, ok := caseInsensitiveMap(oi1.UserDefined).Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
return replicateMetadata
}
// Compare only necessary headers
compareKeys := []string{
@@ -1269,9 +1300,6 @@ func replicateObject(ctx context.Context, ri ReplicateObjectInfo, objectAPI Obje
oi.UserDefined[targetResetHeader(rinfo.Arn)] = rinfo.ResyncTimestamp
}
}
if ri.UserTags != "" {
oi.UserDefined[xhttp.AmzObjectTagging] = ri.UserTags
}
return dsc, nil
},
}
@@ -1689,14 +1717,11 @@ applyAction:
if _, ok := lkMap.Lookup(xhttp.AmzObjectLockRetainUntilDate); ok {
dstOpts.Internal.RetentionTimestamp = objInfo.ModTime
}
if objInfo.UserTags != "" {
dstOpts.Internal.TaggingTimestamp = objInfo.ModTime
}
if tagTmStr, ok := lkMap.Lookup(ReservedMetadataPrefixLower + TaggingTimestamp); ok {
ondiskTimestamp, err := time.Parse(time.RFC3339, tagTmStr)
if err == nil {
dstOpts.Internal.TaggingTimestamp = ondiskTimestamp
}
dstOpts.Internal.TaggingTimestamp, rinfo.Err = replicationTaggingTimestamp(objInfo)
if rinfo.Err != nil {
rinfo.ReplicationStatus = replication.Failed
replLogIf(ctx, fmt.Errorf("invalid tagging timestamp for object %s/%s(%s): %w", bucket, object, objInfo.VersionID, rinfo.Err))
return rinfo
}
if retTmStr, ok := lkMap.Lookup(ReservedMetadataPrefixLower + ObjectLockRetentionTimestamp); ok {
ondiskTimestamp, err := time.Parse(time.RFC3339, retTmStr)
@@ -1916,11 +1941,26 @@ func filterReplicationStatusMetadata(metadata map[string]string) map[string]stri
// DeletedObjectReplicationInfo has info on deleted object
type DeletedObjectReplicationInfo struct {
DeletedObject
Bucket string
EventType string
OpType replication.Type
ResetID string
TargetArn string
Bucket string
EventType string
OpType replication.Type
ResetID string
TargetArn string
RetryCount int
}
// isVersionPurge also recognizes the old marker-shaped purge task. Use the
// operation's state, rather than one target's possibly missing purge entry.
func (di DeletedObjectReplicationInfo) isVersionPurge() bool {
return di.VersionID != "" || di.DeleteMarkerVersionID != "" && !di.VersionPurgeStatus().Empty()
}
// Purge metadata uses COMPLETE; operation statistics and audit use COMPLETED.
func purgeReplicationStatus(status VersionPurgeStatusType) replication.StatusType {
if replication.StatusType(status) == replication.CompletedLegacy {
return replication.Completed
}
return replication.StatusType(status)
}
// ToMRFEntry returns the relevant info needed by MRF
@@ -1930,9 +1970,10 @@ func (di DeletedObjectReplicationInfo) ToMRFEntry() MRFReplicateEntry {
versionID = di.VersionID
}
return MRFReplicateEntry{
Bucket: di.Bucket,
Object: di.ObjectName,
versionID: versionID,
Bucket: di.Bucket,
Object: di.ObjectName,
versionID: versionID,
RetryCount: di.RetryCount,
}
}
@@ -2424,6 +2465,7 @@ func (p *ReplicationPool) queueReplicaDeleteTask(doi DeletedObjectReplicationInf
case <-p.ctx.Done():
case ch <- doi:
default:
doi.RetryCount++
p.queueMRFSave(doi.ToMRFEntry())
p.mu.RLock()
prio := p.priority
@@ -3787,9 +3829,10 @@ func queueReplicationHeal(ctx context.Context, bucket string, oi ObjectInfo, rcf
DeleteMarkerMTime: DeleteMarkerMTime{roi.ModTime},
DeleteMarker: roi.DeleteMarker,
},
Bucket: roi.Bucket,
OpType: replication.HealReplicationType,
EventType: ReplicateHealDelete,
Bucket: roi.Bucket,
OpType: replication.HealReplicationType,
EventType: ReplicateHealDelete,
RetryCount: retryCount,
}
// heal delete marker replication failure or versioned delete replication failure
if roi.ReplicationStatus == replication.Pending ||
@@ -4072,7 +4115,12 @@ func (p *ReplicationPool) queueMRFHeal() error {
VersionID: vID,
})
cancel()
if err != nil {
// A versioned marker lookup returns its metadata with a 405. Only
// accept that error with a real, matching marker identity.
validMarker := isErrMethodNotAllowed(err) && oi.DeleteMarker &&
vID != "" && oi.VersionID == vID && !oi.ModTime.IsZero() &&
oi.Bucket == e.Bucket && oi.Name != "" && oi.Name == decodeDirObject(e.Object)
if err != nil && !validMarker || oi.Name == "" {
continue
}
+1
View File
@@ -445,6 +445,7 @@ func buildServerCtxt(ctx *cli.Context, ctxt *serverCtxt) (err error) {
ctxt.SendBufSize = ctx.Int("send-buf-size")
ctxt.RecvBufSize = ctx.Int("recv-buf-size")
ctxt.IdleTimeout = ctx.Duration("idle-timeout")
ctxt.ReadHeaderTimeout = ctx.Duration("read-header-timeout")
ctxt.UserTimeout = ctx.Duration("conn-user-timeout")
if conf := ctx.String("config"); len(conf) > 0 {
+385
View File
@@ -0,0 +1,385 @@
// Copyright (c) 2026 mr javad seydi and Ruohang Feng
//
// This file is part of Silo Object Storage stack.
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/pgsty/silo-pkg/v3/policy"
)
const multipartScanEntryLimit = 100_000
var (
multipartScanSlots = make(chan struct{}, 2)
errMultipartListingLegacy = errors.New("legacy multipart uploads require a coordinated upgrade and drain")
errMultipartListingIdentity = errors.New("multipart upload identity is invalid")
)
// One budget and admission slot cover the entire request, including all pools.
// A slot is released only after the scan workers have actually stopped.
type multipartScan struct {
ctx context.Context
cancel context.CancelFunc
remaining atomic.Int64
metadataSlots chan struct{}
preflight bool
sets []multipartScanSet
}
type multipartScanSet struct {
Pool int `json:"pool"`
Set int `json:"set"`
Drives int `json:"drives"`
ScannedDrives int `json:"scannedDrives"`
UncoveredDrives []int `json:"uncoveredDrives,omitempty"`
Candidates int `json:"candidates"`
LegacyUploads int `json:"legacyUploads"`
OldestLegacy time.Time `json:"oldestLegacy,omitempty"`
Error string `json:"error,omitempty"`
}
type multipartPreflightReport struct {
Ready bool `json:"ready"`
Complete bool `json:"complete"`
Mode string `json:"mode"`
ScannedEntries int64 `json:"scannedEntries"`
LegacyUploads int `json:"legacyUploads"`
Sets []multipartScanSet `json:"sets"`
}
func startMultipartScan(ctx context.Context, preflight bool) (*multipartScan, error) {
select {
case multipartScanSlots <- struct{}{}:
case <-ctx.Done():
return nil, ctx.Err()
default:
return nil, SlowDown{}
}
ctx, cancel := context.WithTimeout(ctx, 30*time.Second)
s := &multipartScan{ctx: ctx, cancel: cancel, preflight: preflight, metadataSlots: make(chan struct{}, multipartMetadataScanConcurrency)}
s.remaining.Store(multipartScanEntryLimit)
return s, nil
}
func (s *multipartScan) close() {
s.cancel()
<-multipartScanSlots
}
func (s *multipartScan) listDir(disk StorageAPI, bucket, dir string) ([]string, error) {
if err := s.ctx.Err(); err != nil {
return nil, SlowDown{}
}
remaining := s.remaining.Load()
if remaining <= 0 {
return nil, SlowDown{}
}
// Request one extra entry to detect overflow, never a silently partial page.
entries, err := disk.ListDir(s.ctx, bucket, minioMetaMultipartBucket, dir, int(remaining)+1)
if errors.Is(err, errFileNotFound) {
return nil, nil
}
if err != nil {
return nil, err
}
if s.remaining.Add(-int64(len(entries))) < 0 {
return nil, SlowDown{}
}
return entries, nil
}
func (s *multipartScan) listUploadDirs(disk StorageAPI, bucket string) ([]string, error) {
hashDirs, err := s.listDir(disk, bucket, "")
if err != nil {
return nil, err
}
var candidates []string
for _, hashDir := range hashDirs {
if !strings.HasSuffix(hashDir, SlashSeparator) {
continue
}
hashDir = strings.TrimSuffix(hashDir, SlashSeparator)
uploadDirs, err := s.listDir(disk, bucket, hashDir)
if err != nil {
return nil, err
}
for _, uploadDir := range uploadDirs {
if strings.HasSuffix(uploadDir, SlashSeparator) {
candidates = append(candidates, pathJoin(hashDir, strings.TrimSuffix(uploadDir, SlashSeparator)))
}
}
}
return candidates, nil
}
// Identity is immutable at hash/uploadUUID. Check the hash before using even
// a single source drive to exclude another bucket. Missing/bad identities must
// fall back to the full metadata read; they cannot prove absence.
func (er erasureObjects) multipartIdentity(fi FileInfo, shaDir string) (string, string, bool) {
bucket, object := fi.Metadata[multipartMetaBucket], fi.Metadata[multipartMetaObject]
return bucket, object, bucket != "" && object != "" && IsValidBucketName(bucket) &&
IsValidObjectPrefix(object) && er.getMultipartSHADir(bucket, object) == shaDir
}
func (er erasureObjects) readMultipartUploadCandidate(s *multipartScan, bucket, candidate string, source StorageAPI) (MultipartInfo, bool, bool, error) {
if s.ctx.Err() != nil {
return MultipartInfo{}, false, false, SlowDown{}
}
shaDir, uploadUUID, ok := strings.Cut(candidate, SlashSeparator)
if !ok || shaDir == "" || uploadUUID == "" || strings.Contains(uploadUUID, SlashSeparator) {
return MultipartInfo{}, false, false, errMultipartListingIdentity
}
if !s.preflight && bucket != "" && source != nil {
fi, err := source.ReadVersion(s.ctx, bucket, minioMetaMultipartBucket, candidate, "", ReadOptions{})
if err == nil {
storedBucket, _, valid := er.multipartIdentity(fi, shaDir)
if valid && storedBucket != bucket {
return MultipartInfo{}, false, false, nil
}
}
}
if s.ctx.Err() != nil {
return MultipartInfo{}, false, false, SlowDown{}
}
select {
case s.metadataSlots <- struct{}{}:
defer func() { <-s.metadataSlots }()
case <-s.ctx.Done():
return MultipartInfo{}, false, false, SlowDown{}
}
disks := er.getDisks()
metadata, errs := readAllFileInfo(s.ctx, disks, bucket, minioMetaMultipartBucket, candidate, "", false, false)
if s.preflight {
// Readiness is stronger than listing liveness: even one readable legacy
// copy must be drained, and an unreadable copy cannot certify readiness.
var oldest MultipartInfo
legacy := false
_, nativeID := multipartUploadTime(uploadUUID)
for i, err := range errs {
if errors.Is(err, errFileNotFound) || errors.Is(err, errFileVersionNotFound) {
continue
}
if err != nil {
return MultipartInfo{}, false, false, err
}
fi := metadata[i]
storedBucket, storedObject, valid := er.multipartIdentity(fi, shaDir)
if storedBucket != "" && storedObject != "" && !valid {
return MultipartInfo{}, false, false, errMultipartListingIdentity
}
if !valid || !nativeID {
info := multipartUploadInfo(storedBucket, storedObject, uploadUUID, fi.ModTime)
if !legacy || info.Initiated.Before(oldest.Initiated) {
oldest = info
}
legacy = true
}
}
if legacy {
return oldest, false, true, nil
}
}
readQuorum, _, err := objectQuorumFromMeta(s.ctx, metadata, errs, er.defaultParityCount)
if err != nil {
return MultipartInfo{}, false, false, err
}
_, modTime, etag := listOnlineDisks(disks, metadata, errs, readQuorum)
if err := reduceReadQuorumErrs(s.ctx, errs, objectOpIgnoredErrs, readQuorum); err != nil {
return MultipartInfo{}, false, false, err
}
fi, err := pickValidFileInfo(s.ctx, metadata, modTime, etag, readQuorum)
if err != nil {
return MultipartInfo{}, false, false, err
}
storedBucket, storedObject, valid := er.multipartIdentity(fi, shaDir)
info := multipartUploadInfo(storedBucket, storedObject, uploadUUID, fi.ModTime)
if storedBucket == "" || storedObject == "" {
return info, false, true, nil
}
if !valid {
return MultipartInfo{}, false, false, fmt.Errorf("%w: %s", errMultipartListingIdentity, candidate)
}
if _, ok := multipartUploadTime(uploadUUID); !ok {
return info, false, true, nil
}
if bucket != "" && bucket != storedBucket {
return MultipartInfo{}, false, false, nil
}
return info, true, false, nil
}
func (er erasureObjects) scanMultipartUploads(s *multipartScan, bucket string, poolIdx, setIdx int) ([]MultipartInfo, bool, error) {
disks := er.getDisks()
report := multipartScanSet{Pool: poolIdx, Set: setIdx, Drives: er.setDriveCount}
var resultErr error
defer func() {
if resultErr != nil {
report.Error = resultErr.Error()
}
s.sets = append(s.sets, report)
}()
var candidateMu sync.Mutex
candidates := make(map[string]StorageAPI)
errs := make([]error, len(disks))
var wg sync.WaitGroup
for i, disk := range disks {
wg.Add(1)
go func() {
defer wg.Done()
if disk == nil || !disk.IsOnline() {
errs[i] = errDiskNotFound
return
}
paths, err := s.listUploadDirs(disk, bucket)
errs[i] = err
if err != nil {
return
}
candidateMu.Lock()
defer candidateMu.Unlock()
for _, p := range paths {
candidates[p] = disk
}
}()
}
wg.Wait()
for i, err := range errs {
if err == nil {
report.ScannedDrives++
} else {
report.UncoveredDrives = append(report.UncoveredDrives, i)
if _, limited := err.(SlowDown); limited {
resultErr = err
}
}
}
report.Candidates = len(candidates)
// A successful scan must intersect every metadata read quorum, including
// records left with R copies by a partially failed cancellation.
if report.ScannedDrives < er.setDriveCount/2+1 && resultErr == nil {
resultErr = toObjectErr(errErasureReadQuorum, bucket)
}
if resultErr != nil {
return nil, false, resultErr
}
type candidate struct {
path string
source StorageAPI
}
jobs := make(chan candidate)
var mu sync.Mutex
var uploads []MultipartInfo
for range min(16, len(candidates)) {
wg.Add(1)
go func() {
defer wg.Done()
for c := range jobs {
mu.Lock()
stopped := resultErr != nil
mu.Unlock()
if stopped || s.ctx.Err() != nil {
continue
}
upload, found, legacy, err := er.readMultipartUploadCandidate(s, bucket, c.path, c.source)
if errors.Is(err, errFileNotFound) || errors.Is(err, errFileVersionNotFound) {
continue
}
mu.Lock()
switch {
case err != nil && resultErr == nil:
resultErr = err
case legacy:
report.LegacyUploads++
if report.OldestLegacy.IsZero() || upload.Initiated.Before(report.OldestLegacy) {
report.OldestLegacy = upload.Initiated
}
case found && !s.preflight:
uploads = append(uploads, upload)
}
mu.Unlock()
}
}()
}
dispatch:
for p, source := range candidates {
select {
case jobs <- candidate{p, source}:
case <-s.ctx.Done():
break dispatch
}
}
close(jobs)
wg.Wait()
if s.ctx.Err() != nil {
resultErr = SlowDown{}
}
return uploads, report.LegacyUploads != 0, resultErr
}
// multipartPreflight scans all pools, sets and drives, independent of caches.
// A majority suffices for normal listing; upgrade readiness requires every
// drive to have been inspected. The operator must also upgrade all writers.
func (z *erasureServerPools) multipartPreflight(ctx context.Context) (multipartPreflightReport, error) {
s, err := startMultipartScan(ctx, true)
if err != nil {
return multipartPreflightReport{}, err
}
defer s.close()
report := multipartPreflightReport{Complete: true, Mode: "strict"}
if globalAPIConfig.getMultipartListingLegacy() {
report.Mode = "legacy"
}
for p, pool := range z.serverPools {
for i, set := range pool.sets {
_, _, _ = set.scanMultipartUploads(s, "", p, i)
}
}
report.Sets = s.sets
for _, set := range s.sets {
report.LegacyUploads += set.LegacyUploads
if set.Error != "" || set.ScannedDrives != set.Drives {
report.Complete = false
}
}
report.ScannedEntries = multipartScanEntryLimit - s.remaining.Load()
report.Ready = report.Complete && report.LegacyUploads == 0
return report, nil
}
// MultipartPreflightHandler is a read-only storage-admin diagnostic. It never
// accepts an arbitrary deletion path or changes upload lifetime settings.
func (a adminAPIHandlers) MultipartPreflightHandler(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
obj, _ := validateAdminReq(ctx, w, r, policy.StorageInfoAdminAction)
if obj == nil {
return
}
z, ok := obj.(*erasureServerPools)
if !ok {
writeErrorResponseJSON(ctx, w, errorCodes.ToAPIErr(ErrNotImplemented), r.URL)
return
}
report, err := z.multipartPreflight(ctx)
if err != nil {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, err), r.URL)
return
}
b, err := json.Marshal(report)
if err != nil {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, err), r.URL)
return
}
writeSuccessResponseJSON(w, b)
}
+819
View File
@@ -0,0 +1,819 @@
// Copyright (c) 2026 Ruohang Feng
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/base64"
"encoding/json"
"encoding/xml"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"slices"
"strconv"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/minio/minio/internal/config/storageclass"
)
// Check the real handler and storage path, including a successful abort between
// pages. Choose IDs increasing in both initiation time and lexical order so the
// failure does not depend on differing interpretations of S3 marker ordering.
func TestMultipartListingAbortBetweenHTTPPages(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, []string{"ListMultipartUploads", "AbortMultipart"}, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
setMultipartListingTestMode(t, false)
var firstID, secondID string
for attempt := 0; attempt < 32; attempt++ {
one, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
two, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if one.UploadID < two.UploadID {
firstID, secondID = one.UploadID, two.UploadID
break
}
for _, id := range []string{one.UploadID, two.UploadID} {
if err := z.AbortMultipartUpload(t.Context(), bucket, "a", id, ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
}
if firstID == "" {
t.Fatal("could not construct increasing upload IDs")
}
if _, err := z.NewMultipartUpload(t.Context(), bucket, "b", ObjectOptions{}); err != nil {
t.Fatal(err)
}
request := func(method, u string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, u, 0, nil, globalActiveCred.AccessKey, globalActiveCred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
return rec
}
list := func(keyMarker, uploadMarker string, limit int) ListMultipartUploadsResponse {
t.Helper()
rec := request(http.MethodGet, getListMultipartUploadsURLWithParams("", bucket, "", keyMarker, uploadMarker, "", strconv.Itoa(limit)))
if rec.Code != http.StatusOK {
t.Fatalf("list HTTP %d: %s", rec.Code, rec.Body.String())
}
var result ListMultipartUploadsResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &result); err != nil {
t.Fatal(err)
}
return result
}
first := list("", "", 1)
if len(first.Uploads) != 1 || first.Uploads[0].UploadID != firstID || !first.IsTruncated {
t.Fatalf("unexpected first page: %+v", first)
}
rec := request(http.MethodDelete, getAbortMultipartUploadURL("", bucket, "a", firstID))
if rec.Code != http.StatusNoContent {
t.Fatalf("abort HTTP %d: %s", rec.Code, rec.Body.String())
}
rest := list(first.NextKeyMarker, first.NextUploadIDMarker, 10)
var keys []string
for _, upload := range rest.Uploads {
keys = append(keys, upload.Key)
}
t.Logf("GET first page=200; DELETE marker=204; GET next page=200 keys=%v truncated=%v", keys, rest.IsTruncated)
if _, err := z.GetMultipartInfo(t.Context(), bucket, "a", secondID, ObjectOptions{}); err != nil {
t.Fatalf("remaining upload is not valid: %v", err)
}
if len(rest.Uploads) != 2 || rest.Uploads[0].UploadID != secondID {
t.Fatalf("valid remaining upload for key a is missing from continuation: %+v", rest.Uploads)
}
}
type multipartListingFaultDisk struct {
StorageAPI
read func(context.Context, string, string, string, string, ReadOptions) (FileInfo, error)
delete func(context.Context, string, string, DeleteOptions) error
}
func (d multipartListingFaultDisk) ReadVersion(ctx context.Context, original, volume, object, version string, opts ReadOptions) (FileInfo, error) {
if d.read != nil {
return d.read(ctx, original, volume, object, version, opts)
}
return d.StorageAPI.ReadVersion(ctx, original, volume, object, version, opts)
}
func (d multipartListingFaultDisk) Delete(ctx context.Context, volume, object string, opts DeleteOptions) error {
if d.delete != nil {
return d.delete(ctx, volume, object, opts)
}
return d.StorageAPI.Delete(ctx, volume, object, opts)
}
func TestMultipartListingLegacyPreflight(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
const other = "multipart-legacy-other"
if err := z.MakeBucket(t.Context(), other, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
old, err := z.NewMultipartUpload(t.Context(), other, "old", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
fi, metadata, err := set.checkUploadIDExists(t.Context(), other, "old", old.UploadID, true)
if err != nil {
t.Fatal(err)
}
for i := range metadata {
delete(metadata[i].Metadata, multipartMetaBucket)
delete(metadata[i].Metadata, multipartMetaObject)
}
if _, err = writeAllMetadata(t.Context(), set.getDisks(), other, minioMetaMultipartBucket,
set.getUploadIDDir(other, "old", old.UploadID), metadata, fi.WriteQuorum(set.defaultWQuorum())); err != nil {
t.Fatal(err)
}
if _, err = z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{}); err != nil {
t.Fatal(err)
}
z.mpCache.Clear()
t.Run("legacy-upgrade", func(t *testing.T) {
setMultipartListingTestMode(t, true)
exact, err := z.ListMultipartUploads(t.Context(), other, "old", "", "", "", 10)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, exact, "old")
const empty = "multipart-legacy-empty"
if err := z.MakeBucket(t.Context(), empty, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
got, err := z.ListMultipartUploads(t.Context(), empty, "", "", "", "", 10)
if err != nil || len(got.Uploads) != 0 {
t.Fatalf("old upload in another bucket broke legacy listing: %+v %v", got, err)
}
})
_, err = z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 10)
if !errors.Is(err, errMultipartListingLegacy) {
t.Fatalf("old upload in another bucket: %v", err)
}
report, err := z.multipartPreflight(t.Context())
if err != nil || report.Ready || !report.Complete || report.LegacyUploads != 1 {
t.Fatalf("legacy preflight: %+v %v", report, err)
}
if err = z.AbortMultipartUpload(t.Context(), other, "old", old.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
report, err = z.multipartPreflight(t.Context())
if err != nil || !report.Ready || report.LegacyUploads != 0 {
t.Fatalf("drained preflight: %+v %v", report, err)
}
got, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 10)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, got, "a")
}
func TestMultipartListingIdentityFallback(t *testing.T) {
for _, kind := range []string{"old", "corrupt", "read-failure", "wrong-bucket"} {
t.Run(kind, func(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
if _, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{}); err != nil {
t.Fatal(err)
}
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
var first atomic.Bool
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i, d := range disks {
disks[i] = multipartListingFaultDisk{StorageAPI: d, read: func(ctx context.Context, b, v, p, version string, opts ReadOptions) (FileInfo, error) {
fi, err := d.ReadVersion(ctx, b, v, p, version, opts)
if v == minioMetaMultipartBucket && first.CompareAndSwap(false, true) {
if kind == "read-failure" {
return FileInfo{}, errDiskNotFound
}
if kind == "corrupt" {
return FileInfo{}, errFileCorrupt
}
fi.Metadata = cloneMSS(fi.Metadata)
if kind == "old" {
delete(fi.Metadata, multipartMetaBucket)
} else {
fi.Metadata[multipartMetaBucket] = "another-bucket"
}
}
return fi, err
}}
}
return disks
}
got, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 10)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, got, "a")
})
}
}
func TestMultipartAbortPoolsAndRetry(t *testing.T) {
z, bucket := consistencyPools(t)
setMultipartListingTestMode(t, false)
mp, err := z.serverPools[1].NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
emptySet := z.serverPools[0].getHashedSet("a")
original := emptySet.getDisks
t.Cleanup(func() { emptySet.getDisks = original })
emptySet.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i, d := range disks {
disks[i] = multipartListingFaultDisk{StorageAPI: d, read: func(context.Context, string, string, string, string, ReadOptions) (FileInfo, error) {
return FileInfo{}, errDiskNotFound
}}
}
return disks
}
err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{})
if err == nil || toAPIError(t.Context(), err).HTTPStatusCode != 503 {
t.Fatalf("unknown pool must not acknowledge cancellation: %v", err)
}
emptySet.getDisks = original
err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{})
if _, ok := err.(InvalidUploadID); !ok {
t.Fatalf("retry after all pools confirm absence: %v", err)
}
mp, err = z.serverPools[1].NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if _, err = z.serverPools[1].GetMultipartInfo(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err == nil {
t.Fatal("empty first pool hid the actual upload")
}
}
func TestMultipartAbortRetryBelowReadQuorum(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
disks[3] = nil
d := disks[2]
disks[2] = multipartListingFaultDisk{StorageAPI: d, delete: func(ctx context.Context, v, p string, opts DeleteOptions) error {
if err := d.Delete(ctx, v, p, opts); err != nil {
return err
}
return context.DeadlineExceeded // operation finished, acknowledgement lost
}}
return disks
}
if err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err == nil {
t.Fatal("lost acknowledgement must fail")
}
set.getDisks = original
if err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("one remaining metadata copy prevented retry: %v", err)
}
}
func TestMultipartListingMarkerHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, []string{"ListMultipartUploads"}, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
setMultipartListingTestMode(t, false)
for _, tc := range []struct {
key, marker string
status int
}{
{"", "not-base64=", 200},
{"a", "not-base64=", 404},
{"a", base64.RawURLEncoding.EncodeToString([]byte("not-native")), 400},
{"a", multipartListingTestID(time.Unix(100, 0), 1), 200},
} {
u := getListMultipartUploadsURLWithParams("", bucket, "", tc.key, tc.marker, "", "10")
req, err := newTestSignedRequestV4(http.MethodGet, u, 0, nil, globalActiveCred.AccessKey, globalActiveCred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != tc.status {
t.Fatalf("key=%q marker=%q: %d %s", tc.key, tc.marker, rec.Code, rec.Body.String())
}
}
}
func TestMultipartPreflightAdminHTTP(t *testing.T) {
bed, err := prepareAdminErasureTestBed(t.Context())
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { bed.done(); bed.objLayer.Shutdown(context.Background()); removeRoots(bed.erasureDirs) })
const target = "/minio/admin/v3/multipart-preflight"
rec := httptest.NewRecorder()
bed.router.ServeHTTP(rec, httptest.NewRequest(http.MethodGet, target, nil))
if rec.Code != 403 {
t.Fatalf("anonymous preflight: %d", rec.Code)
}
req, err := newTestSignedRequestV4(http.MethodGet, target, 0, nil, globalActiveCred.AccessKey, globalActiveCred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec = httptest.NewRecorder()
bed.router.ServeHTTP(rec, req)
if rec.Code != 200 {
t.Fatalf("admin preflight: %d %s", rec.Code, rec.Body.String())
}
var report multipartPreflightReport
if err = json.Unmarshal(rec.Body.Bytes(), &report); err != nil || !report.Ready || !report.Complete {
t.Fatalf("preflight: %+v %v", report, err)
}
for range cap(multipartScanSlots) {
scan, err := startMultipartScan(t.Context(), false)
if err != nil {
t.Fatal(err)
}
defer scan.close()
}
rec = httptest.NewRecorder()
bed.router.ServeHTTP(rec, req)
if rec.Code != http.StatusServiceUnavailable {
t.Fatalf("busy admin preflight: %d %s", rec.Code, rec.Body.String())
}
}
type multipartLateCreateDisk struct {
StorageAPI
late chan<- func() error
}
func (d multipartLateCreateDisk) WriteMetadata(_ context.Context, original, volume, object string, fi FileInfo) error {
if volume == minioMetaMultipartBucket {
d.late <- func() error { return d.StorageAPI.WriteMetadata(context.Background(), original, volume, object, fi) }
return context.DeadlineExceeded
}
return d.StorageAPI.WriteMetadata(context.Background(), original, volume, object, fi)
}
// This characterization is deliberately NOT an assertion of terminal abort
// correctness. It preserves an executable example of the separately scoped
// late-creation-write limitation. Change the expectation when fencing is added.
func TestMultipartAbortLateCreateBoundary(t *testing.T) {
obj, dirs, err := prepareErasure(t.Context(), 16)
if err != nil {
t.Fatal(err)
}
z := obj.(*erasureServerPools)
setMultipartListingTestMode(t, false)
t.Cleanup(func() { z.Shutdown(context.Background()); removeRoots(dirs) })
saved := globalStorageClass
globalStorageClass.Update(storageclass.Config{Standard: storageclass.StorageClass{Parity: 8}})
t.Cleanup(func() { globalStorageClass.Update(saved) })
const bucket = "multipart-late-create"
if err = z.MakeBucket(t.Context(), bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
set := z.serverPools[0].getHashedSet("a")
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
late := make(chan func() error, 16)
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i := 9; i < 16; i++ {
disks[i] = multipartLateCreateDisk{disks[i], late}
}
return disks
}
mp, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if len(late) != 7 {
t.Fatalf("expected seven timed-out creation writes, got %d", len(late))
}
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i := 2; i < 9; i++ {
disks[i] = nil
}
return disks
}
if err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
for range 7 {
if err = (<-late)(); err != nil {
t.Fatal(err)
}
}
set.getDisks = original
_, metadata, err := set.checkUploadIDExists(t.Context(), bucket, "a", mp.UploadID, true)
if err != nil {
t.Fatalf("late creation boundary changed; review and remove the documented limitation: %v", err)
}
count := 0
for _, fi := range metadata {
if fi.IsValid() {
count++
}
}
if count != 14 {
t.Fatalf("expected fourteen resurrected copies, got %d", count)
}
t.Log("KNOWN UNRESOLVED BOUNDARY: seven delayed creation writes plus seven offline copies restore a writable upload after acknowledged cancellation")
}
type multipartListingCountingDisk struct {
StorageAPI
directoryCalls *atomic.Int64
metadataCalls *atomic.Int64
}
func (d multipartListingCountingDisk) ListDir(ctx context.Context, original, volume, dir string, count int) ([]string, error) {
if volume == minioMetaMultipartBucket {
d.directoryCalls.Add(1)
}
return d.StorageAPI.ListDir(ctx, original, volume, dir, count)
}
func (d multipartListingCountingDisk) ReadVersion(ctx context.Context, original, volume, object, version string, opts ReadOptions) (FileInfo, error) {
if volume == minioMetaMultipartBucket {
d.metadataCalls.Add(1)
}
return d.StorageAPI.ReadVersion(ctx, original, volume, object, version, opts)
}
// Counts storage API calls; this is not a deployment throughput benchmark.
func TestMultipartListingScanCosts(t *testing.T) {
obj, dirs, err := prepareErasure(t.Context(), 4)
if err != nil {
t.Fatal(err)
}
z := obj.(*erasureServerPools)
setMultipartListingTestMode(t, false)
t.Cleanup(func() { z.Shutdown(context.Background()); removeRoots(dirs) })
const bucket, otherBucket = "r9-scan-target", "r9-scan-unrelated"
for _, name := range []string{bucket, otherBucket} {
if err := z.MakeBucket(t.Context(), name, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
}
if _, err := z.NewMultipartUpload(t.Context(), bucket, "dir/target", ObjectOptions{}); err != nil {
t.Fatal(err)
}
var directoryCalls, metadataCalls atomic.Int64
for _, pool := range z.serverPools {
for _, set := range pool.sets {
original := set.getDisks
set.getDisks = func() []StorageAPI {
disks := original()
wrapped := make([]StorageAPI, len(disks))
for i, disk := range disks {
if disk != nil {
wrapped[i] = multipartListingCountingDisk{disk, &directoryCalls, &metadataCalls}
}
}
return wrapped
}
t.Cleanup(func() { set.getDisks = original })
}
}
measure := func(label string) (int64, int64) {
t.Helper()
directoryCalls.Store(0)
metadataCalls.Store(0)
result, err := z.ListMultipartUploads(t.Context(), bucket, "dir/", "", "", "", 1)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, result, "dir/target")
d, m := directoryCalls.Load(), metadataCalls.Load()
t.Logf("%s: max-uploads=1 returned=%d ListDir=%d ReadVersion=%d", label, len(result.Uploads), d, m)
return d, m
}
_, baseline := measure("no unrelated uploads")
for n := 0; n < 32; n++ {
if _, err := z.NewMultipartUpload(t.Context(), otherBucket, fmt.Sprintf("other/%03d", n), ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
_, loaded := measure("32 uploads in another bucket")
_, repeated := measure("same query repeated")
if loaded != baseline+32 || repeated != loaded {
t.Fatalf("unexpected scan accounting: baseline=%d loaded=%d repeated=%d", baseline, loaded, repeated)
}
}
func multipartListingTestID(created time.Time, n int) string {
return base64.RawURLEncoding.EncodeToString([]byte(fmt.Sprintf("00000000-0000-4000-8000-000000000000.00000000-0000-4000-8000-%012xx%d", n, created.UnixNano())))
}
func multipartListingFixture(t *testing.T) (*erasureServerPools, *erasureObjects, string) {
t.Helper()
obj, dirs, err := prepareErasure(t.Context(), 4)
if err != nil {
t.Fatal(err)
}
z := obj.(*erasureServerPools)
t.Cleanup(func() { z.Shutdown(context.Background()); removeRoots(dirs) })
const bucket = "multipart-listing-test"
if err := z.MakeBucket(t.Context(), bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
setMultipartListingTestMode(t, false)
return z, z.serverPools[0].getHashedSet("a"), bucket
}
func TestMultipartListingMissingMarkers(t *testing.T) {
base := time.Unix(100, 0)
sameTime := []string{multipartListingTestID(base.Add(time.Second), 1), multipartListingTestID(base.Add(time.Second), 2), multipartListingTestID(base.Add(time.Second), 3)}
slices.Sort(sameTime)
uploads := []MultipartInfo{
{Bucket: "bucket", Object: "a", UploadID: multipartListingTestID(base, 1), Initiated: base},
{Bucket: "bucket", Object: "a", UploadID: sameTime[1], Initiated: base.Add(time.Second)},
{Bucket: "bucket", Object: "b", UploadID: multipartListingTestID(base, 4), Initiated: base},
}
for _, tc := range []struct {
name string
created time.Time
id int
want []string
}{
{"before", base.Add(-time.Second), 0, []string{"a", "a", "b"}},
{"between", base.Add(time.Second / 2), 2, []string{"a", "b"}},
{"same-time-before", base.Add(time.Second), 2, []string{"a", "b"}},
{"same-time-after", base.Add(time.Second), 4, []string{"b"}},
{"after", base.Add(2 * time.Second), 5, []string{"b"}},
} {
t.Run(tc.name, func(t *testing.T) {
marker := multipartListingTestID(tc.created, tc.id)
if tc.name == "same-time-before" {
marker = sameTime[0]
}
if tc.name == "same-time-after" {
marker = sameTime[2]
}
if err := checkListMultipartArgs(t.Context(), "bucket", "", "a", marker, ""); err != nil {
t.Fatal(err)
}
got := paginateMultipartUploads(uploads, "", "a", marker, "", 10)
if !slices.Equal(multipartUploadKeys(got.Uploads), tc.want) {
t.Fatalf("got %v want %v", multipartUploadKeys(got.Uploads), tc.want)
}
})
}
}
func TestMultipartListingAbortRecovery(t *testing.T) {
for _, offline := range []int{1, 2} {
t.Run(fmt.Sprint(offline), func(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i := 0; i < offline; i++ {
disks[i] = nil
}
return disks
}
err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{})
if offline == 1 && err != nil {
t.Fatal(err)
}
if offline == 2 && (err == nil || toAPIError(t.Context(), err).HTTPStatusCode != 503) {
t.Fatalf("two offline drives must not acknowledge cancellation: %v", err)
}
set.getDisks = original
if offline == 2 {
// Retry a partially completed deletion after the original disks return.
if err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
got, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 100)
if err != nil || len(got.Uploads) != 0 {
t.Fatalf("canceled upload reappeared: %+v %v", got, err)
}
})
}
}
func TestMultipartListingCoverage(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
disks := original()
// Leave a readable two-copy record, simulating a partially failed abort.
for _, d := range disks[:2] {
if err = d.Delete(t.Context(), minioMetaMultipartBucket, set.getUploadIDDir(bucket, "a", mp.UploadID), DeleteOptions{Recursive: true}); err != nil {
t.Fatal(err)
}
}
set.getDisks = func() []StorageAPI { return []StorageAPI{disks[0], disks[1], nil, nil} }
_, err = z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 10)
if err == nil || toAPIError(t.Context(), err).HTTPStatusCode != 503 {
t.Fatalf("incomplete discovery returned success: %v", err)
}
set.getDisks = func() []StorageAPI { return []StorageAPI{nil, disks[1], disks[2], disks[3]} }
got, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 10)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, got, "a")
report, err := z.multipartPreflight(t.Context())
if err != nil || report.Ready || report.Complete || len(report.Sets[0].UncoveredDrives) != 1 {
t.Fatalf("offline upgrade preflight: %+v %v", report, err)
}
}
func TestMultipartListingBudgetAndAdmission(t *testing.T) {
_, set, bucket := multipartListingFixture(t)
if _, err := set.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{}); err != nil {
t.Fatal(err)
}
scan, err := startMultipartScan(t.Context(), false)
if err != nil {
t.Fatal(err)
}
defer scan.close()
scan.remaining.Store(1)
_, _, err = set.scanMultipartUploads(scan, bucket, 0, 0)
var limited SlowDown
if !errors.As(err, &limited) {
t.Fatalf("budget: %v", err)
}
if apiErr := toAPIError(t.Context(), err); apiErr.HTTPStatusCode != http.StatusServiceUnavailable || apiErr.Code != "SlowDown" {
t.Fatalf("budget error mapping: %+v", apiErr)
}
second, err := startMultipartScan(t.Context(), false)
if err != nil {
t.Fatal(err)
}
defer second.close()
if third, err := startMultipartScan(t.Context(), false); err == nil {
third.close()
t.Fatal("third scan admitted")
}
}
func TestMultipartListingAdmissionHTTP(t *testing.T) {
z, _, _ := multipartListingFixture(t)
bucket, router, err := initAPIHandlerTest(t.Context(), z, []string{"ListMultipartUploads"}, MakeBucketOptions{})
if err != nil {
t.Fatal(err)
}
setMultipartListingTestMode(t, false)
for range cap(multipartScanSlots) {
scan, err := startMultipartScan(t.Context(), false)
if err != nil {
t.Fatal(err)
}
defer scan.close()
}
req, err := newTestSignedRequestV4(http.MethodGet,
getListMultipartUploadsURLWithParams("", bucket, "", "", "", "", "1"),
0, nil, globalActiveCred.AccessKey, globalActiveCred.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
var response APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if rec.Code != http.StatusServiceUnavailable || response.Code != "SlowDown" {
t.Fatalf("admission returned %d %s", rec.Code, rec.Body.String())
}
}
func TestMultipartListingPreflightMinorityLegacy(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "old", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
path := set.getUploadIDDir(bucket, "old", mp.UploadID)
disks := set.getDisks()
fi, err := disks[0].ReadVersion(t.Context(), bucket, minioMetaMultipartBucket, path, "", ReadOptions{})
if err != nil {
t.Fatal(err)
}
delete(fi.Metadata, multipartMetaBucket)
delete(fi.Metadata, multipartMetaObject)
if err = disks[0].WriteMetadata(t.Context(), bucket, minioMetaMultipartBucket, path, fi); err != nil {
t.Fatal(err)
}
for _, d := range disks[1:] {
if err = d.Delete(t.Context(), minioMetaMultipartBucket, path, DeleteOptions{Recursive: true}); err != nil {
t.Fatal(err)
}
}
report, err := z.multipartPreflight(t.Context())
if err != nil || report.Ready || !report.Complete || report.LegacyUploads != 1 {
t.Fatalf("minority legacy copy must prevent readiness: %+v %v", report, err)
}
}
func TestMultipartListingCancellationRetainsAdmission(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
for i := range 40 {
if _, err := z.NewMultipartUpload(t.Context(), bucket, fmt.Sprintf("key-%02d", i), ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
original := set.getDisks
t.Cleanup(func() { set.getDisks = original })
entered := make(chan struct{}, 40)
release := make(chan struct{})
var once sync.Once
defer once.Do(func() { close(release) })
var reads atomic.Int32
set.getDisks = func() []StorageAPI {
disks := append([]StorageAPI(nil), original()...)
for i, d := range disks {
disks[i] = multipartListingFaultDisk{StorageAPI: d, read: func(ctx context.Context, b, v, p, version string, opts ReadOptions) (FileInfo, error) {
if v != minioMetaMultipartBucket {
return d.ReadVersion(ctx, b, v, p, version, opts)
}
reads.Add(1)
entered <- struct{}{}
// Model an RPC which does not return immediately on cancellation.
<-release
return FileInfo{}, ctx.Err()
}}
}
return disks
}
second, err := startMultipartScan(t.Context(), false)
if err != nil {
t.Fatal(err)
}
defer second.close()
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
done := make(chan error, 1)
go func() {
_, err := z.ListMultipartUploads(ctx, bucket, "", "", "", "", 10)
done <- err
}()
for range 16 {
select {
case <-entered:
case <-time.After(5 * time.Second):
t.Fatal("identity workers did not start")
}
}
cancel()
if extra, err := startMultipartScan(t.Context(), false); err == nil {
extra.close()
t.Fatal("canceled request released admission while RPCs were still running")
}
start := time.Now()
once.Do(func() { close(release) })
select {
case err := <-done:
if err == nil {
t.Fatal("canceled scan returned a successful partial list")
}
case <-time.After(2 * time.Second):
t.Fatal("scan did not stop after blocked RPCs returned")
}
if got := reads.Load(); got != 16 {
t.Fatalf("scheduled more identity/metadata reads after cancellation: %d", got)
}
t.Logf("16 identity RPCs bounded; admission held until return; cancellation settled in %s", time.Since(start))
}
+331
View File
@@ -0,0 +1,331 @@
// Copyright (c) 2026 Ruohang Feng
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/base64"
"errors"
"fmt"
"strings"
"sync/atomic"
"testing"
"github.com/minio/minio/internal/config/api"
"github.com/minio/minio/internal/config/storageclass"
"github.com/minio/minio/internal/dsync"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
func setMultipartListingTestMode(t *testing.T, legacy bool) {
t.Helper()
globalAPIConfig.mu.Lock()
previous := globalAPIConfig.multipartListingStrict
globalAPIConfig.multipartListingStrict = !legacy
globalAPIConfig.mu.Unlock()
t.Cleanup(func() {
globalAPIConfig.mu.Lock()
globalAPIConfig.multipartListingStrict = previous
globalAPIConfig.mu.Unlock()
})
}
func TestMultipartListingDefaultMode(t *testing.T) {
var uninitialized apiConfig
if !uninitialized.getMultipartListingLegacy() {
t.Fatal("runtime zero value enabled strict mode before config initialization")
}
for _, mode := range []string{"", "legacy", "strict"} {
var local apiConfig
local.init(api.Config{RequestsMax: 1, MultipartListing: mode}, []int{4}, false)
if local.getMultipartListingLegacy() != (mode != "strict") {
t.Fatalf("unexpected mode for %q", mode)
}
}
}
func TestMultipartAbortLegacyAvailability(t *testing.T) {
for _, drives := range []int{4, 16} {
for _, offline := range []int{0, 1, drives / 2} {
t.Run(fmt.Sprintf("drives=%d/offline=%d", drives, offline), func(t *testing.T) {
obj, dirs, err := prepareErasure(t.Context(), drives)
if err != nil {
t.Fatal(err)
}
z := obj.(*erasureServerPools)
t.Cleanup(func() { z.Shutdown(context.Background()); removeRoots(dirs) })
setMultipartListingTestMode(t, true)
previous := globalStorageClass
globalStorageClass.Update(storageclass.Config{Standard: storageclass.StorageClass{Parity: drives / 2}})
t.Cleanup(func() { globalStorageClass.Update(previous) })
const bucket, key = "multipart-legacy-availability", "keep-available"
if err := z.MakeBucket(t.Context(), bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
mp, err := z.NewMultipartUpload(t.Context(), bucket, key, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
set := z.serverPools[0].getHashedSet(key)
original := set.getDisks
disks := original()
set.getDisks = func() []StorageAPI {
visible := append([]StorageAPI(nil), disks...)
clear(visible[:offline])
return visible
}
t.Cleanup(func() { set.getDisks = original })
if err := z.AbortMultipartUpload(t.Context(), bucket, key, mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("released read-quorum availability regressed: %v", err)
}
for _, disk := range disks[offline:] {
_, err := disk.ReadVersion(t.Context(), bucket, minioMetaMultipartBucket, set.getUploadIDDir(bucket, key, mp.UploadID), "", ReadOptions{})
if !errors.Is(err, errFileNotFound) {
t.Fatalf("online replica not cleaned: %v", err)
}
}
})
}
}
}
func TestMultipartAbortLegacyPoolOrder(t *testing.T) {
for _, owner := range []int{0, 1} {
t.Run(fmt.Sprintf("owner=%d", owner), func(t *testing.T) {
z, bucket := consistencyPools(t)
setMultipartListingTestMode(t, true)
mp, err := z.serverPools[owner].NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
emptySet := z.serverPools[1-owner].getHashedSet("a")
original := emptySet.getDisks
t.Cleanup(func() { emptySet.getDisks = original })
emptySet.getDisks = func() []StorageAPI { return make([]StorageAPI, emptySet.setDriveCount) }
err = z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{})
if owner == 0 && err != nil {
t.Fatalf("unrelated later pool blocked legacy cancellation: %v", err)
}
if owner == 1 {
if err == nil || toAPIError(t.Context(), err).HTTPStatusCode != 503 {
t.Fatalf("unknown earlier pool must retain released error behavior: %v", err)
}
if _, err := z.serverPools[owner].GetMultipartInfo(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("later pool visited despite earlier error: %v", err)
}
emptySet.getDisks = original
if err := z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
})
}
}
func TestMultipartAbortMinorityLegacyCleanup(t *testing.T) {
for _, failDelete := range []bool{false, true} {
t.Run(fmt.Sprintf("delete-fails=%v", failDelete), func(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "old", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
path := set.getUploadIDDir(bucket, "old", mp.UploadID)
original := set.getDisks
disks := original()
fi, err := disks[0].ReadVersion(t.Context(), bucket, minioMetaMultipartBucket, path, "", ReadOptions{})
if err != nil {
t.Fatal(err)
}
delete(fi.Metadata, multipartMetaBucket)
delete(fi.Metadata, multipartMetaObject)
if err := disks[0].WriteMetadata(t.Context(), bucket, minioMetaMultipartBucket, path, fi); err != nil {
t.Fatal(err)
}
for _, d := range disks[1:] {
if err := d.Delete(t.Context(), minioMetaMultipartBucket, path, DeleteOptions{Recursive: true}); err != nil {
t.Fatal(err)
}
}
report, err := z.multipartPreflight(t.Context())
if err != nil || report.Ready || report.LegacyUploads != 1 {
t.Fatalf("invalid legacy fixture: %+v %v", report, err)
}
if failDelete {
set.getDisks = func() []StorageAPI {
wrapped := append([]StorageAPI(nil), disks...)
wrapped[0] = multipartListingFaultDisk{StorageAPI: disks[0], delete: func(context.Context, string, string, DeleteOptions) error { return errFaultyDisk }}
return wrapped
}
t.Cleanup(func() { set.getDisks = original })
err := z.AbortMultipartUpload(t.Context(), bucket, "old", mp.UploadID, ObjectOptions{})
if err == nil || toAPIError(t.Context(), err).HTTPStatusCode != 503 {
t.Fatalf("empty drives masked failed deletion of the observed replica: %v", err)
}
if _, present := z.mpCache.Load(mp.UploadID); !present {
t.Fatal("failed cancellation evicted upload from cache")
}
set.getDisks = original
}
if err := z.AbortMultipartUpload(t.Context(), bucket, "old", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
report, err = z.multipartPreflight(t.Context())
if err != nil || !report.Ready || report.LegacyUploads != 0 {
t.Fatalf("acknowledged cleanup left online legacy remnants: %+v %v", report, err)
}
if _, present := z.mpCache.Load(mp.UploadID); present {
t.Fatal("successful cancellation retained cache entry")
}
})
}
}
func TestMultipartAbortWrongTargetKeepsCache(t *testing.T) {
for _, legacy := range []bool{true, false} {
t.Run(fmt.Sprintf("legacy=%v", legacy), func(t *testing.T) {
z, _, bucket := multipartListingFixture(t)
setMultipartListingTestMode(t, legacy)
const other = "multipart-other-bucket"
if err := z.MakeBucket(t.Context(), other, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
mp, err := z.NewMultipartUpload(t.Context(), bucket, "valid-key", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
for _, target := range [][2]string{{bucket, "wrong-key"}, {other, "valid-key"}} {
err := z.AbortMultipartUpload(t.Context(), target[0], target[1], mp.UploadID, ObjectOptions{})
var invalid InvalidUploadID
if !errors.As(err, &invalid) {
t.Fatalf("wrong target result: %v", err)
}
if _, present := z.mpCache.Load(mp.UploadID); !present {
t.Fatal("wrong-target abort evicted another upload")
}
if _, err := z.GetMultipartInfo(t.Context(), bucket, "valid-key", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("wrong-target abort harmed valid upload: %v", err)
}
}
if err := z.AbortMultipartUpload(t.Context(), bucket, "valid-key", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
var invalid InvalidUploadID
if err := z.AbortMultipartUpload(t.Context(), bucket, "valid-key", mp.UploadID, ObjectOptions{}); !errors.As(err, &invalid) {
t.Fatalf("repeat abort: %v", err)
}
})
}
}
func TestMultipartAbortRejectsUnsafeID(t *testing.T) {
for _, legacy := range []bool{true, false} {
t.Run(fmt.Sprintf("legacy=%v", legacy), func(t *testing.T) {
z, _, bucket := multipartListingFixture(t)
setMultipartListingTestMode(t, legacy)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "keep", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
for _, suffix := range []string{".", "..", "../other", "a/b", "a\\b", ""} {
id := base64.RawURLEncoding.EncodeToString([]byte("deployment." + suffix))
err := z.AbortMultipartUpload(t.Context(), bucket, "keep", id, ObjectOptions{})
var invalid InvalidUploadID
if !errors.As(err, &invalid) {
t.Fatalf("unsafe ID accepted: suffix=%q err=%v", suffix, err)
}
if _, err := z.GetMultipartInfo(t.Context(), bucket, "keep", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("valid upload harmed by rejected ID: %v", err)
}
}
})
}
}
func TestMultipartAbortPeerNotification(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
tg, err := grid.SetupTestGrid(2)
if err != nil {
t.Fatal(err)
}
t.Cleanup(tg.Cleanup)
var notifications atomic.Int32
if err := cleanupUploadIDCacheMetaRPC.Register(tg.Managers[1], func(*grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
notifications.Add(1)
return grid.NoPayload{}, nil
}); err != nil {
t.Fatal(err)
}
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
if err != nil {
t.Fatal(err)
}
previous := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) },
}}}
t.Cleanup(func() { globalNotificationSys = previous })
// Use the real distributed lock implementation: Unlock cancels its derived
// context. Local locks do not, and would hide a broken notification context.
oldMutex, oldLockers := set.nsMutex, set.getLockers
lockers := []dsync.NetLocker{newLocker(), newLocker()}
set.nsMutex = newNSLock(true)
set.getLockers = func() ([]dsync.NetLocker, string) { return lockers, "multipart-test" }
t.Cleanup(func() { set.nsMutex, set.getLockers = oldMutex, oldLockers })
for _, legacy := range []bool{true, false} {
t.Run(fmt.Sprintf("legacy=%v", legacy), func(t *testing.T) {
setMultipartListingTestMode(t, legacy)
for range 4 {
mp, err := z.NewMultipartUpload(t.Context(), bucket, "valid", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
before := notifications.Load()
var invalid InvalidUploadID
if err := z.AbortMultipartUpload(t.Context(), bucket, "wrong", mp.UploadID, ObjectOptions{}); !errors.As(err, &invalid) {
t.Fatalf("wrong target: %v", err)
}
if notifications.Load() != before {
t.Fatal("wrong-target abort sent a destructive peer notification")
}
if err := z.AbortMultipartUpload(t.Context(), bucket, "valid", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
if notifications.Load() != before+1 {
t.Fatal("successful abort lost peer notification after unlocking")
}
}
})
}
}
func TestMultipartAbortStrictMajority(t *testing.T) {
z, set, bucket := multipartListingFixture(t)
mp, err := z.NewMultipartUpload(t.Context(), bucket, "a", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original := set.getDisks
disks := original()
set.getDisks = func() []StorageAPI {
wrapped := append([]StorageAPI(nil), disks...)
wrapped[0] = multipartListingFaultDisk{StorageAPI: disks[0], delete: func(context.Context, string, string, DeleteOptions) error { return errFaultyDisk }}
return wrapped
}
t.Cleanup(func() { set.getDisks = original })
if err := z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatalf("ordinary majority cancellation was tightened: %v", err)
}
// One failed deletion is tolerated in the ordinary majority path. Retrying
// after recovery must also clean the now-minority remnant.
if _, err := disks[0].ReadVersion(t.Context(), bucket, minioMetaMultipartBucket, set.getUploadIDDir(bucket, "a", mp.UploadID), "", ReadOptions{}); err != nil {
t.Fatalf("fault fixture did not leave a remnant: %v", err)
}
set.getDisks = original
if err := z.AbortMultipartUpload(t.Context(), bucket, "a", mp.UploadID, ObjectOptions{}); err != nil {
t.Fatal(err)
}
}
+271 -35
View File
@@ -31,6 +31,7 @@ import (
"sync"
"time"
"github.com/google/uuid"
"github.com/klauspost/readahead"
"github.com/minio/minio-go/v7/pkg/set"
"github.com/minio/minio/internal/config/storageclass"
@@ -44,6 +45,15 @@ import (
"github.com/pgsty/silo-pkg/v3/sync/errgroup"
)
const (
multipartMetaBucket = ReservedMetadataPrefixLower + "multipart-v1-bucket"
multipartMetaObject = ReservedMetadataPrefixLower + "multipart-v1-object"
// ponytail: keep scan concurrency fixed until the multipart-list benchmark
// establishes a better adaptive limit.
multipartMetadataScanConcurrency = 4
)
func (er erasureObjects) getUploadIDDir(bucket, object, uploadID string) string {
uploadUUID := uploadID
uploadBytes, err := base64.RawURLEncoding.DecodeString(uploadID)
@@ -251,14 +261,152 @@ func (er erasureObjects) cleanupStaleUploadsOnDisk(ctx context.Context, disk Sto
})
}
// ListMultipartUploads - lists all the pending multipart
// uploads for a particular object in a bucket.
//
// Implements minimal S3 compatible ListMultipartUploads API. We do
// not support prefix based listing, this is a deliberate attempt
// towards simplification of multipart APIs.
// The resulting ListMultipartsInfo structure is unmarshalled directly as XML.
func (er erasureObjects) ListMultipartUploads(ctx context.Context, bucket, object, keyMarker, uploadIDMarker, delimiter string, maxUploads int) (result ListMultipartsInfo, err error) {
func multipartUploadInfo(bucket, object, uploadUUID string, fallback time.Time) MultipartInfo {
initiated := fallback
if parsed, ok := multipartUploadTime(uploadUUID); ok {
initiated = parsed
}
return MultipartInfo{
Bucket: bucket,
Object: object,
UploadID: base64.RawURLEncoding.EncodeToString(fmt.Appendf(nil, "%s.%s", globalDeploymentID(), uploadUUID)),
Initiated: initiated,
}
}
// multipartUploadTime is shared by stored records and continuation markers.
// The time is part of the immutable upload ID, so a removed marker still
// identifies the same ordering boundary.
func multipartUploadTime(uploadUUID string) (time.Time, bool) {
if len(uploadUUID) < 38 || uploadUUID[36] != 'x' {
return time.Time{}, false
}
if _, err := uuid.Parse(uploadUUID[:36]); err != nil {
return time.Time{}, false
}
ns, err := strconv.ParseInt(uploadUUID[37:], 10, 64)
if err != nil || ns <= 0 || strconv.FormatInt(ns, 10) != uploadUUID[37:] {
return time.Time{}, false
}
return time.Unix(0, ns), true
}
func multipartMarkerTime(uploadID string) (time.Time, bool) {
b, err := base64.RawURLEncoding.DecodeString(uploadID)
if err != nil {
return time.Time{}, false
}
_, uploadUUID, ok := strings.Cut(string(b), ".")
if !ok {
return time.Time{}, false
}
return multipartUploadTime(uploadUUID)
}
type multipartListEntry struct {
upload *MultipartInfo
commonPrefix string
}
// paginateMultipartUploads applies the S3 ordering, prefix, delimiter, marker,
// and page rules exactly once after all pools and sets have been merged.
func paginateMultipartUploads(uploads []MultipartInfo, prefix, keyMarker, uploadIDMarker, delimiter string, maxUploads int) ListMultipartsInfo {
if maxUploads > maxUploadsList {
maxUploads = maxUploadsList
}
result := ListMultipartsInfo{
MaxUploads: maxUploads,
KeyMarker: keyMarker,
UploadIDMarker: uploadIDMarker,
Prefix: prefix,
Delimiter: delimiter,
}
deduplicated := make([]MultipartInfo, 0, len(uploads))
seenUploads := make(map[string]struct{}, len(uploads))
for _, upload := range uploads {
identity := upload.Bucket + "\x00" + upload.Object + "\x00" + upload.UploadID
if _, ok := seenUploads[identity]; ok {
continue
}
seenUploads[identity] = struct{}{}
deduplicated = append(deduplicated, upload)
}
sort.Slice(deduplicated, func(i, j int) bool {
if deduplicated[i].Object != deduplicated[j].Object {
return deduplicated[i].Object < deduplicated[j].Object
}
if !deduplicated[i].Initiated.Equal(deduplicated[j].Initiated) {
return deduplicated[i].Initiated.Before(deduplicated[j].Initiated)
}
return deduplicated[i].UploadID < deduplicated[j].UploadID
})
markerTime, _ := multipartMarkerTime(uploadIDMarker)
seenPrefixes := make(map[string]struct{})
entries := make([]multipartListEntry, 0, len(deduplicated))
for i := range deduplicated {
upload := &deduplicated[i]
if !strings.HasPrefix(upload.Object, prefix) {
continue
}
if keyMarker != "" {
switch strings.Compare(upload.Object, keyMarker) {
case -1:
continue
case 0:
if uploadIDMarker == "" || upload.Initiated.Before(markerTime) ||
(upload.Initiated.Equal(markerTime) && upload.UploadID <= uploadIDMarker) {
continue
}
}
}
if delimiter != "" {
remainder := strings.TrimPrefix(upload.Object, prefix)
if i := strings.Index(remainder, delimiter); i >= 0 {
commonPrefix := prefix + remainder[:i+len(delimiter)]
if keyMarker != "" && commonPrefix <= keyMarker {
continue
}
if _, ok := seenPrefixes[commonPrefix]; ok {
continue
}
seenPrefixes[commonPrefix] = struct{}{}
entries = append(entries, multipartListEntry{commonPrefix: commonPrefix})
continue
}
}
entries = append(entries, multipartListEntry{upload: upload})
}
if maxUploads <= 0 {
return result
}
pageSize := min(maxUploads, len(entries))
for _, entry := range entries[:pageSize] {
if entry.upload != nil {
result.Uploads = append(result.Uploads, *entry.upload)
continue
}
result.CommonPrefixes = append(result.CommonPrefixes, entry.commonPrefix)
}
result.IsTruncated = pageSize < len(entries)
if result.IsTruncated && pageSize > 0 {
last := entries[pageSize-1]
if last.upload != nil {
result.NextKeyMarker = last.upload.Object
result.NextUploadIDMarker = last.upload.UploadID
} else {
result.NextKeyMarker = last.commonPrefix
}
}
return result
}
// listMultipartUploadsExact preserves the hashed exact-object lookup used by
// multipart write placement and by rolling-upgrade legacy mode.
func (er erasureObjects) listMultipartUploadsExact(ctx context.Context, bucket, object, keyMarker, uploadIDMarker, delimiter string, maxUploads int) (result ListMultipartsInfo, err error) {
auditObjectErasureSet(ctx, "ListMultipartUploads", object, &er)
result.MaxUploads = maxUploads
@@ -311,20 +459,16 @@ func (er erasureObjects) ListMultipartUploads(ctx context.Context, bucket, objec
if populatedUploadIDs.Contains(uploadID) {
continue
}
// If present, use time stored in ID.
startTime := time.Now()
if split := strings.Split(uploadID, "x"); len(split) == 2 {
t, err := strconv.ParseInt(split[1], 10, 64)
if err == nil {
startTime = time.Unix(0, t)
var fallback time.Time
if _, ok := multipartUploadTime(uploadID); !ok {
fi, err := disk.ReadVersion(ctx, bucket, minioMetaMultipartBucket,
pathJoin(er.getMultipartSHADir(bucket, object), uploadID), "", ReadOptions{})
if err != nil {
return result, toObjectErr(err, bucket, object)
}
fallback = fi.ModTime
}
uploads = append(uploads, MultipartInfo{
Bucket: bucket,
Object: object,
UploadID: base64.RawURLEncoding.EncodeToString(fmt.Appendf(nil, "%s.%s", globalDeploymentID(), uploadID)),
Initiated: startTime,
})
uploads = append(uploads, multipartUploadInfo(bucket, object, uploadID, fallback))
populatedUploadIDs.Add(uploadID)
}
@@ -365,6 +509,25 @@ func (er erasureObjects) ListMultipartUploads(ctx context.Context, bucket, objec
return result, nil
}
func (er erasureObjects) ListMultipartUploads(ctx context.Context, bucket, prefix, keyMarker, uploadIDMarker, delimiter string, maxUploads int) (ListMultipartsInfo, error) {
if err := checkListMultipartArgs(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter); err != nil {
return ListMultipartsInfo{}, err
}
scan, err := startMultipartScan(ctx, false)
if err != nil {
return ListMultipartsInfo{}, err
}
defer scan.close()
uploads, legacy, err := er.scanMultipartUploads(scan, bucket, 0, 0)
if err != nil {
return ListMultipartsInfo{}, err
}
if legacy {
return ListMultipartsInfo{}, errMultipartListingLegacy
}
return paginateMultipartUploads(uploads, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads), nil
}
// newMultipartUpload - wrapper for initializing a new multipart
// request; returns a unique upload id.
//
@@ -402,6 +565,8 @@ func (er erasureObjects) newMultipartUpload(ctx context.Context, bucket string,
}
userDefined := cloneMSS(opts.UserDefined)
userDefined[multipartMetaBucket] = bucket
userDefined[multipartMetaObject] = object
if opts.PreserveETag != "" {
userDefined["etag"] = opts.PreserveETag
}
@@ -1459,6 +1624,8 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
// Remove superfluous internal headers.
delete(fi.Metadata, hash.MinIOMultipartChecksum)
delete(fi.Metadata, hash.MinIOMultipartChecksumType)
delete(fi.Metadata, multipartMetaBucket)
delete(fi.Metadata, multipartMetaObject)
// Save the final object size and modtime.
fi.Size = objectSize
@@ -1586,22 +1753,91 @@ func (er erasureObjects) CompleteMultipartUpload(ctx context.Context, bucket str
return fi.ToObjectInfo(bucket, object, opts.Versioned || opts.VersionSuspended), nil
}
// AbortMultipartUpload - aborts an ongoing multipart operation
// signified by the input uploadID. This is an atomic operation
// doesn't require clients to initiate multiple such requests.
//
// All parts are purged from all disks and reference to the uploadID
// would be removed from the system, rollback is not possible on this
// operation.
func (er erasureObjects) AbortMultipartUpload(ctx context.Context, bucket, object, uploadID string, opts ObjectOptions) (err error) {
// abortMultipartUpload retains read-quorum validation and best-effort cleanup
// in legacy mode. Strict mode requires majority deletion acknowledgements and
// permits retrying remnants below read quorum. Neither mode fences creation
// writes still executing after a storage timeout.
func (er erasureObjects) abortMultipartUpload(ctx context.Context, bucket, object, uploadID string, opts ObjectOptions, legacy bool) (bool, error) {
if !opts.NoAuditLog {
auditObjectErasureSet(ctx, "AbortMultipartUpload", object, &er)
}
// Cleanup all uploaded parts.
defer er.deleteAll(ctx, minioMetaMultipartBucket, er.getUploadIDDir(bucket, object, uploadID))
// Validates if upload ID exists.
_, _, err = er.checkUploadIDExists(ctx, bucket, object, uploadID, false)
return toObjectErr(err, bucket, object, uploadID)
b, err := base64.RawURLEncoding.DecodeString(uploadID)
if err != nil {
return false, MalformedUploadID{UploadID: uploadID}
}
_, internalID, ok := strings.Cut(string(b), ".")
if !ok || internalID == "" || internalID == "." || internalID == ".." || strings.ContainsAny(internalID, "/\\") {
return false, InvalidUploadID{Bucket: bucket, Object: object, UploadID: uploadID}
}
if legacy {
// Keep the released read-quorum and best-effort cleanup behavior.
// The upload ID safety check above applies to both modes.
defer er.deleteAll(ctx, minioMetaMultipartBucket, er.getUploadIDDir(bucket, object, uploadID))
_, _, err := er.checkUploadIDExists(ctx, bucket, object, uploadID, false)
err = toObjectErr(err, bucket, object, uploadID)
if _, absent := err.(InvalidUploadID); absent {
return false, nil
}
return err == nil, err
}
disks := er.getDisks()
uploadPath := er.getUploadIDDir(bucket, object, uploadID)
_, errs := readAllFileInfo(ctx, disks, bucket, minioMetaMultipartBucket, uploadPath, "", false, false)
quorum := er.setDriveCount/2 + 1
found, absent := false, 0
observed := make([]bool, len(disks))
for i, err := range errs {
switch {
case err == nil, errors.Is(err, errFileCorrupt):
found = true
observed[i] = true
case errors.Is(err, errFileNotFound), errors.Is(err, errFileVersionNotFound):
absent++
}
}
if absent >= quorum && !found {
return false, nil
}
if !found {
return false, toObjectErr(errErasureReadQuorum, bucket, object, uploadID)
}
g := errgroup.WithNErrs(len(disks))
for i, disk := range disks {
g.Go(func() error {
if disk == nil {
return errDiskNotFound
}
err := disk.Delete(ctx, minioMetaMultipartBucket, uploadPath, DeleteOptions{Recursive: true, Immediate: false})
if errors.Is(err, errFileNotFound) || errors.Is(err, errFileVersionNotFound) {
return nil
}
return err
}, i)
}
deleteErrs := g.Wait()
if err := reduceWriteQuorumErrs(ctx, deleteErrs, nil, quorum); err != nil {
return true, toObjectErr(err, bucket, object, uploadID)
}
if absent >= quorum {
// Missing disks must not mask a failed cleanup of known remnants.
for i, found := range observed {
if found && deleteErrs[i] != nil {
return true, toObjectErr(errErasureWriteQuorum, bucket, object, uploadID)
}
}
}
return true, nil
}
// AbortMultipartUpload cancels an upload using the configured mode. Offline
// part data may still need stale-upload cleanup after its drives return.
func (er erasureObjects) AbortMultipartUpload(ctx context.Context, bucket, object, uploadID string, opts ObjectOptions) (err error) {
found, err := er.abortMultipartUpload(ctx, bucket, object, uploadID, opts, globalAPIConfig.getMultipartListingLegacy())
if err != nil {
return err
}
if !found {
return InvalidUploadID{Bucket: bucket, Object: object, UploadID: uploadID}
}
return nil
}
+4
View File
@@ -2327,7 +2327,11 @@ func (er erasureObjects) PutObjectTags(ctx context.Context, bucket, object strin
fi.Metadata[xhttp.AmzObjectTagging] = tags
fi.ReplicationState = opts.PutReplicationState()
stamp := monotonicTaggingTimestamp(opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp])
maps.Copy(fi.Metadata, opts.UserDefined)
if stamp != "" {
fi.Metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
if err = er.updateObjectMeta(ctx, bucket, object, fi, onlineDisks); err != nil {
return ObjectInfo{}, toObjectErr(err, bucket, object)
+15
View File
@@ -242,6 +242,21 @@ func reconcileStoredObjectTags(metadata map[string]string, storedTags, storedTim
}
}
// Local tagging mutations must advance the revision they overwrite, even when
// a request's clock or lock acquisition order is behind the stored revision.
// Replica writes use reconcileStoredObjectTags instead of minting a revision.
func monotonicTaggingTimestamp(incoming, stored string) string {
requested, err := time.Parse(time.RFC3339Nano, incoming)
if err != nil {
return incoming
}
current, err := time.Parse(time.RFC3339Nano, stored)
if err != nil || requested.After(current) {
return incoming
}
return current.Add(time.Nanosecond).UTC().Format(time.RFC3339Nano)
}
// A restored version still owns its tier reference even while IsRemote is
// false. Only the last copy of a reference may schedule its contents for GC.
func sharesTierObject(oi ObjectInfo, copies []PoolObjInfo) bool {
@@ -0,0 +1,721 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"crypto/md5"
"encoding/base64"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// Change allocation capacity only; metadata and data use real fixture disks.
// Restoring the adapters also lets encrypted fixtures move the write target
// between requests without changing the production allocation policy.
type conditionalPutCapacityDisk struct {
StorageAPI
full bool
}
func (d conditionalPutCapacityDisk) DiskInfo(ctx context.Context, opts DiskInfoOptions) (DiskInfo, error) {
info, err := d.StorageAPI.DiskInfo(ctx, opts)
info.Total, info.Used = 1<<40, 0
if d.full {
info.Used = info.Total - (1 << 20)
}
info.Free = info.Total - info.Used
return info, err
}
// Keep getDisks immutable while background IAM and storage readers use it.
// GetDisks takes this same mutex when it copies the backing disk list.
func conditionalPutSwapDisks(pool *erasureSets, object string, wrap func(StorageAPI) StorageAPI) func() {
setIndex := pool.getHashedSet(object).setIndex
pool.erasureDisksMu.Lock()
previous := pool.erasureDisks[setIndex]
disks := append([]StorageAPI(nil), previous...)
for i, disk := range disks {
disks[i] = wrap(disk)
}
pool.erasureDisks[setIndex] = disks
pool.erasureDisksMu.Unlock()
return func() {
pool.erasureDisksMu.Lock()
pool.erasureDisks[setIndex] = previous
pool.erasureDisksMu.Unlock()
}
}
func conditionalPutPool(t *testing.T, z *erasureServerPools, object string, target int) func() {
t.Helper()
var restore []func()
for i, pool := range z.serverPools {
restore = append(restore, conditionalPutSwapDisks(pool, object, func(disk StorageAPI) StorageAPI {
return conditionalPutCapacityDisk{StorageAPI: disk, full: i != target}
}))
}
return func() {
for _, fn := range restore {
fn()
}
}
}
func conditionalPutBucket(t *testing.T, z *erasureServerPools, mode string) (string, http.Handler) {
t.Helper()
bucket, router, err := initAPIHandlerTest(t.Context(), z, nil, MakeBucketOptions{VersioningEnabled: mode != "unversioned"})
if err != nil {
t.Fatal(err)
}
if mode == "suspended" {
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig,
[]byte(`<VersioningConfiguration><Status>Suspended</Status></VersioningConfiguration>`)); err != nil {
t.Fatal(err)
}
}
return bucket, router
}
func TestPoolsConditionalPutHTTP(t *testing.T) {
for _, mode := range []string{"unversioned", "versioned", "suspended"} {
t.Run(mode, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, mode)
for target := range 2 {
for _, tc := range []struct {
name string
oldCopy, localCurrent bool
condition string
status int
}{
{"stale-etag-accepted", true, false, "old", http.StatusPreconditionFailed},
{"current-etag-rejected", true, false, "current", http.StatusOK},
{"create-only-overwrites-other-pool", false, false, "none", http.StatusPreconditionFailed},
{"current-etag-missing-in-write-pool", false, false, "current", http.StatusOK},
{"control-current-in-write-pool", true, true, "current", http.StatusOK},
{"control-stale-etag-rejected", true, true, "old", http.StatusPreconditionFailed},
{"none-match-current-etag", true, false, "none-current", http.StatusPreconditionFailed},
{"none-match-old-etag", true, false, "none-old", http.StatusOK},
} {
t.Run(fmt.Sprintf("target=%d/%s", target, tc.name), func(t *testing.T) {
object := fmt.Sprintf("%d-%s", target, tc.name)
currentPool := 1 - target
if tc.localCurrent {
currentPool = target
}
opts := ObjectOptions{Versioned: mode == "versioned", VersionSuspended: mode == "suspended", MTime: UTCNow().Add(-time.Hour)}
oldETag := "absent-old"
if tc.oldCopy {
oldETag = putConsistencyObject(t, z, bucket, object, 1-currentPool, "old", opts).ETag
}
opts.MTime = UTCNow().Add(-time.Minute)
current := putConsistencyObject(t, z, bucket, object, currentPool, "current", opts)
defer conditionalPutPool(t, z, object, target)()
idx, err := z.getWritePoolIdx(t.Context(), bucket, object, 11, false)
if err != nil || idx != target {
t.Fatalf("allocation target=%d: idx=%d err=%v", target, idx, err)
}
headers := map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", current.ETag)}
if tc.condition == "old" {
headers[xhttp.IfMatch] = fmt.Sprintf("%q", oldETag)
}
if tc.condition == "none" {
headers = map[string]string{xhttp.IfNoneMatch: "*"}
}
if tc.condition == "none-current" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", current.ETag)}
}
if tc.condition == "none-old" {
headers = map[string]string{xhttp.IfNoneMatch: fmt.Sprintf("%q", oldETag)}
}
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if before.Code != http.StatusOK || before.Body.String() != "current" || multipartConditionResponseETag(before) != current.ETag {
t.Fatalf("invalid current object: %d %q %v", before.Code, before.Body.String(), before.Header())
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != tc.status {
t.Errorf("PUT status=%d want=%d: %s", put.Code, tc.status, put.Body.String())
}
if put.Code == http.StatusPreconditionFailed {
multipartConditionError(t, put, "PreconditionFailed")
if multipartConditionResponseETag(put) != current.ETag || put.Header().Get(xhttp.LastModified) != before.Header().Get(xhttp.LastModified) {
t.Errorf("412 headers do not describe current object: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
wantBody, wantETag := "current", current.ETag
if tc.status == http.StatusOK {
wantBody, wantETag = "replacement", fmt.Sprintf("%x", md5.Sum([]byte("replacement")))
}
t.Logf("PUT %d; GET %d bytes=%q ETag=%s", put.Code, get.Code, get.Body.String(), multipartConditionResponseETag(get))
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Errorf("GET=%d bytes=%q ETag=%s; want %q %s", get.Code, get.Body.String(), multipartConditionResponseETag(get), wantBody, wantETag)
}
})
}
}
})
}
}
func TestPoolsConditionalPutHTTPAbsence(t *testing.T) {
for _, state := range []string{"missing", "uuid-marker", "null-marker"} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("%s/match=%t", state, match), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "versioned"
if state == "null-marker" {
mode = "suspended"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "absent-key"
if state != "missing" {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
vid := mustGetUUID()
if state == "null-marker" {
vid = nullVersionID
}
if _, err := z.serverPools[1].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: vid, DeleteMarker: true, MTime: UTCNow().Add(-time.Minute)}); err != nil {
t.Fatal(err)
}
}
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
want := http.StatusOK
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
want = http.StatusNotFound
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "new", headers)
if put.Code != want {
t.Fatalf("PUT %d want %d: %s", put.Code, want, put.Body.String())
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if match {
multipartConditionError(t, put, "NoSuchKey")
if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT exposed data: %d %s", get.Code, get.Body.String())
}
} else if get.Code != http.StatusOK || get.Body.String() != "new" || multipartConditionResponseETag(get) != multipartConditionResponseETag(put) {
t.Fatalf("successful create GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutUnreadable(t *testing.T) {
for faultPool := range 2 {
for _, present := range []bool{false, true} {
for _, match := range []bool{false, true} {
t.Run(fmt.Sprintf("fault-pool=%d/present=%t/match=%t", faultPool, present, match), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "unreadable-key"
if present {
putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{MTime: UTCNow().Add(-time.Hour)})
putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
}
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[faultPool], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
headers = map[string]string{xhttp.IfMatch: "*"}
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != http.StatusServiceUnavailable {
t.Errorf("unverified PUT must fail: %d %s", put.Code, put.Body.String())
}
called := 0
_, err := z.PutObject(t.Context(), bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), ObjectOptions{HasIfMatch: match, CheckPrecondFn: func(ObjectInfo) bool { called++; return false }})
if !isErrReadQuorum(err) || called != 0 {
t.Errorf("lookup error=%v callback calls=%d", err, called)
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if present {
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != fmt.Sprintf("%x", md5.Sum([]byte("current"))) {
t.Fatalf("failed PUT changed object: %d %q", get.Code, get.Body.String())
}
} else if get.Code != http.StatusNotFound {
t.Fatalf("failed PUT created object: %d %q", get.Code, get.Body.String())
}
})
}
}
}
}
func TestPoolsConditionalPutEncryptedETag(t *testing.T) {
for _, kind := range []string{"SSE-C", "SSE-S3", "SSE-KMS"} {
t.Run(kind, func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
oldKMS, oldTLS := GlobalKMS, globalIsTLS
GlobalKMS, globalIsTLS = kms.NewStub("conditional-put-key"), true
defer func() { GlobalKMS, globalIsTLS = oldKMS, oldTLS }()
headers := map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}
readHeaders := map[string]string{}
if kind == "SSE-C" {
key := bytes.Repeat([]byte{0x42}, 32)
digest := md5.Sum(key)
headers = map[string]string{
xhttp.AmzServerSideEncryptionCustomerAlgorithm: xhttp.AmzEncryptionAES,
xhttp.AmzServerSideEncryptionCustomerKey: base64.StdEncoding.EncodeToString(key),
xhttp.AmzServerSideEncryptionCustomerKeyMD5: base64.StdEncoding.EncodeToString(digest[:]),
}
readHeaders = maps.Clone(headers)
}
if kind == "SSE-KMS" {
headers[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionKMS
headers[xhttp.AmzServerSideEncryptionKmsID] = "conditional-put-key"
}
object := "encrypted-key"
url := getPutObjectURL("", bucket, object)
restore := conditionalPutPool(t, z, object, 0)
old := multipartConditionRequest(t, router, http.MethodPut, url, "old", headers)
restore()
restore = conditionalPutPool(t, z, object, 1)
current := multipartConditionRequest(t, router, http.MethodPut, url, "current", headers)
restore()
if old.Code != http.StatusOK || current.Code != http.StatusOK {
t.Fatalf("encrypted setup: %d %s / %d %s", old.Code, old.Body.String(), current.Code, current.Body.String())
}
defer conditionalPutPool(t, z, object, 0)()
for _, stale := range []bool{true, false} {
h := maps.Clone(headers)
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(current))
wantStatus, wantBody, wantETag := http.StatusOK, "replacement", ""
if stale {
h[xhttp.IfMatch] = fmt.Sprintf("%q", multipartConditionResponseETag(old))
wantStatus, wantBody, wantETag = http.StatusPreconditionFailed, "current", multipartConditionResponseETag(current)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", h)
if put.Code != wantStatus {
t.Fatalf("encrypted condition: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
if !stale {
wantETag = multipartConditionResponseETag(put)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", readHeaders)
if get.Code != http.StatusOK || get.Body.String() != wantBody || multipartConditionResponseETag(get) != wantETag {
t.Fatalf("encrypted GET: %d %q %v", get.Code, get.Body.String(), get.Header())
}
}
})
}
}
func TestPoolsConditionalPutVersionSelection(t *testing.T) {
for _, kind := range []string{"public-version", "replica", "replica-preserve-etag", "movement", "no-lock", "tie", "draining"} {
t.Run(kind, func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "version-selection"
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
opts := ObjectOptions{Versioned: true, VersionID: addressed.VersionID, HasIfMatch: true}
want := current
switch kind {
case "replica":
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "replica-preserve-etag":
opts.PreserveETag = addressed.ETag
opts.ReplicaLockReconcile, opts.ReplicationRequest = true, true
want = addressed
case "movement":
opts.DataMovement, opts.SrcPoolIdx = true, 1
want = addressed
case "tie":
want = putConsistencyObject(t, z, bucket, object, 0, "tie-winner", ObjectOptions{Versioned: true, MTime: current.ModTime})
case "draining":
z.poolMetaMutex.Lock()
z.poolMeta.Pools[1].Decommission = &PoolDecommissionInfo{}
z.poolMetaMutex.Unlock()
}
defer conditionalPutPool(t, z, object, 0)()
ctx := t.Context()
if kind == "no-lock" {
lk := z.NewNSLock(bucket, object)
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
if err != nil {
t.Fatal(err)
}
defer lk.Unlock(lkctx)
ctx, opts.NoLock = lkctx.Context(), true
}
called := 0
opts.UserDefined = make(map[string]string)
opts.CheckPrecondFn = func(oi ObjectInfo) bool {
called++
if oi.ETag != want.ETag || oi.VersionID != want.VersionID {
t.Errorf("comparison ETag/version=%s/%s want %s/%s", oi.ETag, oi.VersionID, want.ETag, want.VersionID)
}
return oi.ETag != want.ETag
}
oi, err := z.PutObject(ctx, bucket, object, mustGetPutObjReader(t, strings.NewReader("replacement"), 11, "", ""), opts)
if err != nil || called != 1 {
t.Fatalf("PUT err=%v callback calls=%d", err, called)
}
if oi.VersionID != addressed.VersionID {
t.Fatalf("destination version changed: %s", oi.VersionID)
}
if opts.PreserveETag != "" && oi.ETag != opts.PreserveETag {
t.Fatalf("PreserveETag changed: %s", oi.ETag)
}
})
}
}
func TestPoolsConditionalPutReplicaDuplicateHTTP(t *testing.T) {
for _, null := range []bool{false, true} {
t.Run(fmt.Sprintf("null=%t", null), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "replica-duplicate"
vid := mustGetUUID()
if null {
vid = nullVersionID
}
addressed := putConsistencyObject(t, z, bucket, object, 1, "addressed", ObjectOptions{Versioned: true, VersionID: vid, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: addressed.ETag,
xhttp.MinIOSourceMTime: addressed.ModTime.Format(time.RFC3339Nano),
}
url := getPutObjectURL("", bucket, object)
put := multipartConditionRequest(t, router, http.MethodPut, url+"?versionId="+vid, "addressed", headers)
if put.Code != http.StatusPreconditionFailed {
t.Fatalf("replica duplicate: %d %s", put.Code, put.Body.String())
}
multipartConditionError(t, put, "PreconditionFailed")
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "current" || multipartConditionResponseETag(get) != current.ETag {
t.Fatalf("duplicate changed current object: %d %q", get.Code, get.Body.String())
}
get = multipartConditionRequest(t, router, http.MethodGet, url+"?versionId="+vid, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "addressed" || multipartConditionResponseETag(get) != addressed.ETag {
t.Fatalf("duplicate changed addressed version: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutConcurrentHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "unversioned")
for _, match := range []bool{false, true} {
for iteration := range 5 {
t.Run(fmt.Sprintf("match=%t/iteration=%d", match, iteration), func(t *testing.T) {
object := fmt.Sprintf("concurrent-%t-%d", match, iteration)
headers := map[string]string{xhttp.IfNoneMatch: "*"}
if match {
oi := putConsistencyObject(t, z, bucket, object, 1, "old", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
headers = map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", oi.ETag)}
}
defer conditionalPutPool(t, z, object, 0)()
start := make(chan struct{})
results := make(chan *httptest.ResponseRecorder, 2)
url := getPutObjectURL("", bucket, object)
for i := range 2 {
body := fmt.Sprintf("writer-%d", i)
req, err := newTestSignedRequestV4(http.MethodPut, url, int64(len(body)), strings.NewReader(body), globalActiveCred.AccessKey, globalActiveCred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
go func() {
<-start
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
results <- rec
}()
}
close(start)
success, failed, winnerETag := 0, 0, ""
for range 2 {
result := <-results
switch result.Code {
case http.StatusOK:
success++
winnerETag = multipartConditionResponseETag(result)
case http.StatusPreconditionFailed:
failed++
multipartConditionError(t, result, "PreconditionFailed")
default:
t.Errorf("unexpected PUT %d: %s", result.Code, result.Body.String())
}
}
if success != 1 || failed != 1 {
t.Fatalf("success=%d precondition failures=%d", success, failed)
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || multipartConditionResponseETag(get) != winnerETag || fmt.Sprintf("%x", md5.Sum(get.Body.Bytes())) != winnerETag {
t.Fatalf("winner lost: %d %q %v", get.Code, get.Body.String(), get.Header())
}
})
}
}
}
func TestPoolsConditionalPutSerializesMutation(t *testing.T) {
for _, deletion := range []bool{false, true} {
t.Run(fmt.Sprintf("delete=%t", deletion), func(t *testing.T) {
z, bucket := consistencyPools(t)
object := "conditional-mutation"
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second)
defer cancel()
gate := &consistencyGateReader{Reader: strings.NewReader("replacement"), entered: make(chan struct{}), resume: make(chan struct{})}
release := func() { gate.release.Do(func() { close(gate.resume) }) }
defer release()
reader := mustGetPutObjReader(t, gate, 11, "", "")
written := make(chan error, 1)
go func() {
_, err := z.PutObject(ctx, bucket, object, reader, ObjectOptions{HasIfMatch: true, CheckPrecondFn: func(oi ObjectInfo) bool { return oi.ETag != current.ETag }})
written <- err
}()
select {
case <-gate.entered:
case err := <-written:
t.Fatalf("PUT failed before body read: %v", err)
case <-ctx.Done():
t.Fatal(ctx.Err())
}
mutated := make(chan error, 1)
go func() {
var err error
if deletion {
_, err = z.DeleteObject(ctx, bucket, object, ObjectOptions{})
} else {
_, err = z.PutObjectMetadata(ctx, bucket, object, ObjectOptions{EvalMetadataFn: func(oi *ObjectInfo, _ error) (ReplicateDecision, error) {
oi.UserDefined["x-amz-meta-after-put"] = "present"
return ReplicateDecision{}, nil
}})
}
mutated <- err
}()
select {
case err := <-mutated:
t.Fatalf("mutation escaped PUT lock: %v", err)
case <-time.After(100 * time.Millisecond):
}
release()
if err := <-written; err != nil {
t.Fatal(err)
}
if err := <-mutated; err != nil {
t.Fatal(err)
}
oi, err := z.GetObjectInfo(ctx, bucket, object, ObjectOptions{})
if deletion {
if !isErrObjectNotFound(err) {
t.Fatalf("delete lost: %v", err)
}
} else if err != nil || oi.UserDefined["x-amz-meta-after-put"] != "present" || oi.ETag != fmt.Sprintf("%x", md5.Sum([]byte("replacement"))) {
t.Fatalf("metadata/PUT lost: %+v %v", oi, err)
}
})
}
}
func TestSinglePoolConditionalPutHTTP(t *testing.T) {
ctx, cancel := context.WithCancel(t.Context())
obj, dirs, err := prepareErasure16(ctx)
if err != nil {
cancel()
t.Fatal(err)
}
z := obj.(*erasureServerPools)
t.Cleanup(func() { cancel(); z.Shutdown(context.Background()); removeRoots(dirs) })
if !z.SinglePool() {
t.Fatal("fixture is not a single pool")
}
bucket, router := conditionalPutBucket(t, z, "unversioned")
object := "single-pool-condition"
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
for _, tc := range []struct {
body, match, none string
status int
}{
{"missing", "*", "", http.StatusNotFound},
{"first", "", "*", http.StatusOK},
{"blocked", "", "*", http.StatusPreconditionFailed},
{"blocked", "stale", "", http.StatusPreconditionFailed},
{"second", fmt.Sprintf("%x", md5.Sum([]byte("first"))), "", http.StatusOK},
} {
rec := multipartConditionRequest(t, router, http.MethodPut, url, tc.body, map[string]string{xhttp.IfMatch: tc.match, xhttp.IfNoneMatch: tc.none})
if rec.Code != tc.status {
t.Fatalf("PUT %d want %d: %s", rec.Code, tc.status, rec.Body.String())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "second" {
t.Fatalf("single-pool GET: %d %q", get.Code, get.Body.String())
}
}
// An internal replica callback without an addressed version is not a public
// condition. Preserve its availability when another pool is unreadable. An
// addressed replica already requires all pools for lock/tag reconciliation.
func TestPoolsConditionalPutReplicaAvailability(t *testing.T) {
for _, addressed := range []bool{false, true} {
t.Run(fmt.Sprintf("addressed=%t", addressed), func(t *testing.T) {
z, _ := consistencyPools(t)
mode := "unversioned"
if addressed {
mode = "versioned"
}
bucket, router := conditionalPutBucket(t, z, mode)
object := "replica-availability"
oi := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: addressed, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
restoreFault := conditionalPutSwapDisks(z.serverPools[1], object, func(disk StorageAPI) StorageAPI {
return consistencyReadFaultDisk{StorageAPI: disk, bucket: bucket, object: object}
})
defer restoreFault()
headers := map[string]string{
xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA",
xhttp.MinIOSourceETag: fmt.Sprintf("%x", md5.Sum([]byte("replacement"))),
}
url := getPutObjectURL("", bucket, object)
wantStatus, wantBody := http.StatusOK, "replacement"
if addressed {
url += "?versionId=" + oi.VersionID
wantStatus, wantBody = http.StatusServiceUnavailable, "old"
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", headers)
if put.Code != wantStatus {
t.Fatalf("replica PUT: %d want %d: %s", put.Code, wantStatus, put.Body.String())
}
restoreFault()
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != wantBody {
t.Fatalf("replica GET: %d %q", get.Code, get.Body.String())
}
})
}
}
func TestPoolsConditionalPutDestinationVersionHTTP(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "client-version"
old := putConsistencyObject(t, z, bucket, object, 0, "old", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Hour)})
current := putConsistencyObject(t, z, bucket, object, 1, "current", ObjectOptions{Versioned: true, MTime: UTCNow().Add(-time.Minute)})
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object) + "?versionId=" + old.VersionID
for _, stale := range []bool{true, false} {
tag, status := current.ETag, http.StatusOK
if stale {
tag, status = old.ETag, http.StatusPreconditionFailed
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfMatch: fmt.Sprintf("%q", tag)})
if put.Code != status {
t.Fatalf("version-addressed public PUT: %d want %d: %s", put.Code, status, put.Body.String())
}
if got := put.Header()[xhttp.AmzVersionID]; !stale && (len(got) != 1 || got[0] != old.VersionID) {
t.Fatalf("write version changed: %v", put.Header())
}
}
get := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "replacement" {
t.Fatalf("addressed GET: %d %q", get.Code, get.Body.String())
}
}
func TestPoolsConditionalPutDeleteMarkerTie(t *testing.T) {
for markerPool := range 2 {
t.Run(fmt.Sprintf("marker-pool=%d", markerPool), func(t *testing.T) {
z, _ := consistencyPools(t)
bucket, router := conditionalPutBucket(t, z, "versioned")
object := "marker-tie"
mtime := UTCNow().Add(-time.Minute)
putConsistencyObject(t, z, bucket, object, 1-markerPool, "live", ObjectOptions{Versioned: true, MTime: mtime})
if _, err := z.serverPools[markerPool].DeleteObject(t.Context(), bucket, object, ObjectOptions{Versioned: true, VersionID: mustGetUUID(), DeleteMarker: true, MTime: mtime}); err != nil {
t.Fatal(err)
}
defer conditionalPutPool(t, z, object, 0)()
url := getPutObjectURL("", bucket, object)
before := multipartConditionRequest(t, router, http.MethodGet, url, "", nil)
want := http.StatusPreconditionFailed
if markerPool == 0 {
want = http.StatusOK
if before.Code != http.StatusNotFound {
t.Fatalf("GET tie: %d", before.Code)
}
} else if before.Code != http.StatusOK {
t.Fatalf("GET tie: %d", before.Code)
}
put := multipartConditionRequest(t, router, http.MethodPut, url, "replacement", map[string]string{xhttp.IfNoneMatch: "*"})
if put.Code != want {
t.Fatalf("PUT tie: %d want %d: %s", put.Code, want, put.Body.String())
}
})
}
}
// Overlap fixture changes with the real IAM Walk reader instead of relying on
// its periodic refresh timer to expose an unsynchronized disk-adapter swap.
func TestPoolsConditionalPutFixtureConcurrentIAM(t *testing.T) {
z, _ := consistencyPools(t)
conditionalPutBucket(t, z, "unversioned")
ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second)
defer cancel()
started, finished := make(chan struct{}), make(chan error, 1)
iam := globalIAMSys
go func() {
close(started)
for range 20 {
if err := iam.Load(ctx, false); err != nil {
finished <- err
return
}
}
finished <- nil
}()
<-started
for i := range 5000 {
restore := conditionalPutPool(t, z, "fixture-concurrent-iam", i%2)
restore()
}
if err := <-finished; err != nil {
t.Fatal(err)
}
}
+103 -17
View File
@@ -1149,6 +1149,40 @@ func (z *erasureServerPools) PutObject(ctx context.Context, bucket string, objec
}
opts.NoLock = true
// Public write conditions compare the logical current object while the
// pools-layer write lock is held. The destination selected by capacity may
// be empty or stale, and draining pools can still hold the current object.
// Replica callbacks retain their existing addressed-version semantics and
// metadata reconciliation at the set layer.
if opts.CheckPrecondFn != nil && !opts.ReplicationRequest &&
!opts.ReplicaLockReconcile && !opts.DataMovement {
copies, lerr := z.objectPoolInfos(ctx, bucket, object, ObjectOptions{
VersionID: "", // Compare the current object, not the write's version.
Versioned: opts.Versioned,
VersionSuspended: opts.VersionSuspended,
NoAuditLog: true,
})
var latest ObjectInfo
if lerr == nil {
latest = copies[0].ObjInfo
if latest.DeleteMarker {
lerr = toObjectErr(errFileNotFound, bucket, object)
}
}
// An unreadable pool may hold the newest object; it is not absence.
if lerr != nil && !isErrObjectNotFound(lerr) && !isErrVersionNotFound(lerr) {
return ObjectInfo{}, lerr
}
if lerr == nil && opts.CheckPrecondFn(latest) {
return ObjectInfo{}, PreConditionFailed{}
}
if lerr != nil && opts.HasIfMatch {
return ObjectInfo{}, lerr
}
// Do not repeat an accepted condition against the destination's copy.
opts.CheckPrecondFn = nil
}
idx, err := z.getWritePoolIdx(ctx, bucket, object, data.Size(), true)
if err != nil {
return ObjectInfo{}, err
@@ -1857,7 +1891,41 @@ func (z *erasureServerPools) ListMultipartUploads(ctx context.Context, bucket, p
if err := checkListMultipartArgs(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter); err != nil {
return ListMultipartsInfo{}, err
}
if _, err := z.GetBucketInfo(ctx, bucket, BucketOptions{}); err != nil {
return ListMultipartsInfo{}, toObjectErr(err, bucket)
}
if globalAPIConfig.getMultipartListingLegacy() {
return z.listMultipartUploadsLegacy(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
}
scan, err := startMultipartScan(ctx, false)
if err != nil {
return ListMultipartsInfo{}, err
}
defer scan.close()
var uploads []MultipartInfo
var keyless bool
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) {
continue
}
poolUploads, poolKeyless, err := pool.scanMultipartUploads(scan, bucket, idx)
if err != nil {
return ListMultipartsInfo{}, err
}
uploads = append(uploads, poolUploads...)
keyless = keyless || poolKeyless
}
// The old format cannot be enumerated authoritatively. Migration mode is
// explicit: another bucket must never silently change this API's semantics.
if keyless {
return ListMultipartsInfo{}, errMultipartListingLegacy
}
return paginateMultipartUploads(uploads, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads), nil
}
func (z *erasureServerPools) listMultipartUploadsLegacy(ctx context.Context, bucket, prefix, keyMarker, uploadIDMarker, delimiter string, maxUploads int) (ListMultipartsInfo, error) {
poolResult := ListMultipartsInfo{}
poolResult.MaxUploads = maxUploads
poolResult.KeyMarker = keyMarker
@@ -1883,15 +1951,14 @@ func (z *erasureServerPools) ListMultipartUploads(ctx context.Context, bucket, p
}
if z.SinglePool() {
return z.serverPools[0].ListMultipartUploads(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
return z.serverPools[0].getHashedSet(prefix).listMultipartUploadsExact(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
}
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) {
continue
}
result, err := pool.ListMultipartUploads(ctx, bucket, prefix, keyMarker, uploadIDMarker,
delimiter, maxUploads)
result, err := pool.getHashedSet(prefix).listMultipartUploadsExact(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
if err != nil {
return result, err
}
@@ -1927,7 +1994,7 @@ func (z *erasureServerPools) NewMultipartUpload(ctx context.Context, bucket, obj
continue
}
result, err := pool.ListMultipartUploads(ctx, bucket, object, "", "", "", maxUploadsList)
result, err := pool.listMultipartUploadsExact(ctx, bucket, object)
if err != nil {
return nil, err
}
@@ -2089,13 +2156,18 @@ func (z *erasureServerPools) AbortMultipartUpload(ctx context.Context, bucket, o
if err := checkAbortMultipartArgs(ctx, bucket, object, uploadID); err != nil {
return err
}
if _, err := z.GetBucketInfo(ctx, bucket, BucketOptions{}); err != nil {
return toObjectErr(err, bucket)
}
defer func() {
// Unlock cancels the derived lock context before this notification runs.
// Keep the request context so successful cancellation reaches peer caches.
defer func(ctx context.Context) {
if err == nil {
z.mpCache.Delete(uploadID)
globalNotificationSys.DeleteUploadID(ctx, uploadID)
}
}()
}(ctx)
lk := z.NewNSLock(bucket, pathJoin(object, uploadID))
lkctx, err := lk.GetLock(ctx, globalOperationTimeout)
@@ -2105,23 +2177,28 @@ func (z *erasureServerPools) AbortMultipartUpload(ctx context.Context, bucket, o
ctx = lkctx.Context()
defer lk.Unlock(lkctx)
if z.SinglePool() {
return z.serverPools[0].AbortMultipartUpload(ctx, bucket, object, uploadID, opts)
}
legacy := globalAPIConfig.getMultipartListingLegacy()
found := false
var firstErr error
for idx, pool := range z.serverPools {
if z.IsSuspended(idx) {
continue
}
err := pool.AbortMultipartUpload(ctx, bucket, object, uploadID, opts)
if err == nil {
return nil
poolFound, err := pool.getHashedSet(object).abortMultipartUpload(ctx, bucket, object, uploadID, opts, legacy)
if legacy && (poolFound || err != nil) {
// Match the released first-matching-pool behavior.
return err
}
if _, ok := err.(InvalidUploadID); ok {
// upload id not found move to next pool
continue
found = found || poolFound
if err != nil && firstErr == nil {
firstErr = err
}
return err
}
if firstErr != nil {
return firstErr
}
if found {
return nil
}
return InvalidUploadID{
Bucket: bucket,
@@ -3051,6 +3128,15 @@ func (z *erasureServerPools) PutObjectTags(ctx context.Context, bucket, object s
if err != nil {
return ObjectInfo{}, err
}
// Ordinary reads and replication can return any owning pool. Persist one
// revision beyond all copies, so the returned value and every pool agree.
if stamp := opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]; stamp != "" {
for _, copy := range copies {
stamp = monotonicTaggingTimestamp(stamp, copy.ObjInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp])
}
opts.UserDefined = cloneMSS(opts.UserDefined)
opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = stamp
}
opts.NoLock = true
opts.VersionID = copies[0].ObjInfo.VersionID
if opts.VersionID == "" {
+34 -4
View File
@@ -880,10 +880,40 @@ func (s *erasureSets) CopyObject(ctx context.Context, srcBucket, srcObject, dstB
}
func (s *erasureSets) ListMultipartUploads(ctx context.Context, bucket, prefix, keyMarker, uploadIDMarker, delimiter string, maxUploads int) (result ListMultipartsInfo, err error) {
// In list multipart uploads we are going to treat input prefix as the object,
// this means that we are not supporting directory navigation.
set := s.getHashedSet(prefix)
return set.ListMultipartUploads(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads)
if err := checkListMultipartArgs(ctx, bucket, prefix, keyMarker, uploadIDMarker, delimiter); err != nil {
return ListMultipartsInfo{}, err
}
scan, err := startMultipartScan(ctx, false)
if err != nil {
return ListMultipartsInfo{}, err
}
defer scan.close()
uploads, legacy, err := s.scanMultipartUploads(scan, bucket, 0)
if err != nil {
return ListMultipartsInfo{}, err
}
if legacy {
return ListMultipartsInfo{}, errMultipartListingLegacy
}
return paginateMultipartUploads(uploads, prefix, keyMarker, uploadIDMarker, delimiter, maxUploads), nil
}
func (s *erasureSets) scanMultipartUploads(scan *multipartScan, bucket string, poolIdx int) ([]MultipartInfo, bool, error) {
var uploads []MultipartInfo
var keyless bool
for i, set := range s.sets {
setUploads, setKeyless, err := set.scanMultipartUploads(scan, bucket, poolIdx, i)
if err != nil {
return nil, false, err
}
uploads = append(uploads, setUploads...)
keyless = keyless || setKeyless
}
return uploads, keyless, nil
}
func (s *erasureSets) listMultipartUploadsExact(ctx context.Context, bucket, object string) (ListMultipartsInfo, error) {
return s.getHashedSet(object).listMultipartUploadsExact(ctx, bucket, object, "", "", "", maxUploadsList)
}
// Initiate a new multipart upload on a hashedSet based on object name.
+8
View File
@@ -50,6 +50,7 @@ type apiConfig struct {
transitionWorkers int
staleUploadsExpiry time.Duration
multipartListingStrict bool
staleUploadsCleanupInterval time.Duration
deleteCleanupInterval time.Duration
enableODirect bool
@@ -181,6 +182,7 @@ func (t *apiConfig) init(cfg api.Config, setDriveCounts []int, legacy bool) {
t.transitionWorkers = cfg.TransitionWorkers
t.staleUploadsExpiry = cfg.StaleUploadsExpiry
t.multipartListingStrict = cfg.MultipartListing == "strict"
t.deleteCleanupInterval = cfg.DeleteCleanupInterval
t.enableODirect = cfg.EnableODirect
t.gzipObjects = cfg.GzipObjects
@@ -206,6 +208,12 @@ func (t *apiConfig) odirectEnabled() bool {
return t.enableODirect
}
func (t *apiConfig) getMultipartListingLegacy() bool {
t.mu.RLock()
defer t.mu.RUnlock()
return !t.multipartListingStrict
}
func (t *apiConfig) shouldGzipObjects() bool {
t.mu.RLock()
defer t.mu.RUnlock()
+33 -23
View File
@@ -246,16 +246,6 @@ func extractMetadata(ctx context.Context, mimesHeader ...textproto.MIMEHeader) (
// extractMetadata extracts metadata from map values.
func extractMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, false)
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
return extractMetadataFromMimeWithReplication(ctx, v, m, true)
}
func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIMEHeader, m map[string]string, allowReplication bool) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
@@ -267,18 +257,14 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
nv[http.CanonicalHeaderKey(k)] = kv
}
// Save all supported headers.
// Save ordinary object metadata. Replication-only headers are restored only
// after the request has been validated as a trusted replication write.
for _, supportedHeader := range supportedHeaders {
value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]
if ok {
if v, ok := replicationToInternalHeaders[supportedHeader]; ok {
if !allowReplication {
continue
}
m[v] = strings.Join(value, ",")
} else {
m[supportedHeader] = strings.Join(value, ",")
}
if _, ok := replicationToInternalHeaders[supportedHeader]; ok {
continue
}
if value, ok := nv[http.CanonicalHeaderKey(supportedHeader)]; ok {
m[supportedHeader] = strings.Join(value, ",")
}
}
@@ -287,8 +273,7 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
if !stringsHasPrefixFold(key, prefix) {
continue
}
value, ok := nv[http.CanonicalHeaderKey(key)]
if ok {
if value, ok := nv[http.CanonicalHeaderKey(key)]; ok {
m[key] = strings.Join(value, ",")
break
}
@@ -297,6 +282,31 @@ func extractMetadataFromMimeWithReplication(ctx context.Context, v textproto.MIM
return nil
}
// extractReplicationMetadataFromMime restores replication-only metadata after the
// caller has validated that the request is a trusted replication write.
func extractReplicationMetadataFromMime(ctx context.Context, v textproto.MIMEHeader, m map[string]string) error {
if v == nil {
bugLogIf(ctx, errInvalidArgument)
return errInvalidArgument
}
nv := make(textproto.MIMEHeader, len(v))
for k, kv := range v {
// Canonicalize all headers, to remove any duplicates.
nv[http.CanonicalHeaderKey(k)] = kv
}
// Ordinary object metadata belongs to the caller. Re-extracting it would
// undo normalization (such as removing aws-chunked) or copy an outer
// Snowball archive's metadata onto its individual entries.
for header, internalHeader := range replicationToInternalHeaders {
if value, ok := nv[http.CanonicalHeaderKey(header)]; ok {
m[internalHeader] = strings.Join(value, ",")
}
}
return nil
}
// Returns access credentials in the request Authorization header.
func getReqAccessCred(r *http.Request, region string) (cred auth.Credentials) {
cred, _, _ = getReqAccessKeyV4(r, region, serviceS3)
+12 -1
View File
@@ -254,6 +254,9 @@ func TestExtractMetadataFromRequestKeepsQueryCompatibility(t *testing.T) {
func TestExtractReplicationMetadataHeaders(t *testing.T) {
header := http.Header{
"Content-Type": []string{"application/wasm"},
"Content-Encoding": []string{"aws-chunked"},
"X-Amz-Meta-Source": []string{"client"},
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key": []string{"sealed-key"},
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm": []string{"DAREv2-HMAC-SHA256"},
"X-Minio-Replication-Server-Side-Encryption-Iv": []string{"iv"},
@@ -262,12 +265,17 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
ReplicationSsecChecksumHeader: []string{"checksum"},
}
metadata := make(map[string]string)
metadata := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
}
if err := extractReplicationMetadataFromMime(t.Context(), textproto.MIMEHeader(header), metadata); err != nil {
t.Fatalf("failed to extract replication metadata: %v", err)
}
expected := map[string]string{
"content-type": "application/wasm",
"x-amz-meta-source": "client",
"X-Minio-Internal-Server-Side-Encryption-Sealed-Key": "sealed-key",
"X-Minio-Internal-Server-Side-Encryption-Seal-Algorithm": "DAREv2-HMAC-SHA256",
"X-Minio-Internal-Server-Side-Encryption-Iv": "iv",
@@ -279,6 +287,9 @@ func TestExtractReplicationMetadataHeaders(t *testing.T) {
if !reflect.DeepEqual(metadata, expected) {
t.Fatalf("unexpected replication metadata: expected %#v, got %#v", expected, metadata)
}
if _, ok := metadata["content-encoding"]; ok {
t.Fatalf("replication metadata restored transport content-encoding: %#v", metadata)
}
}
func TestGetCopyObjectMetadataFromHeaderReplication(t *testing.T) {
+111
View File
@@ -0,0 +1,111 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"errors"
"testing"
"time"
"github.com/minio/minio/internal/auth"
)
func TestIAMCredentialRetention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
secret, err := getTokenSigningKey()
mustIAM(t, err)
parent := "external-idp-parent"
credential := func(exp time.Time) auth.Credentials {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": exp.Unix(), parentClaim: parent}, secret)
mustIAM(t, err)
cred.ParentUser = parent
return cred
}
// Disablement of an external identity must include cached STS,
// which are kept separately from regular and service accounts.
cred := credential(UTCNow().Add(time.Hour))
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
mustIAM(t, err)
mustIAM(t, sys.store.DeleteUsers(ctx, []string{parent}))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
mustIAM(t, err)
if !r.Deleted || !r.ExpiresAt.Equal(cred.Expiration.Add(globalMaxSkewTime)) || r.Credentials.SessionToken != "" || r.Credentials.SecretKey != "" {
t.Fatal("early STS revocation lost its retention boundary or retained a secret")
}
if _, ok := sys.store.GetUser(cred.AccessKey); ok {
t.Fatal("external disablement left the STS cache live")
}
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), cred.AccessKey, cred, "")
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("same revoked token was reissued by replay: %v", err)
}
var mp MappedPolicy
err = sys.store.loadIAMConfig(ctx, &mp, getMappedPolicyPath(cred.AccessKey, stsUser, false))
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("random STS key produced a permanent mapping: %v", err)
}
// Seed genuinely expired immutable tokens, as an ordinary startup
// loader sees them. Natural expiry leaves no permanent tombstone.
expired := credential(UTCNow().Add(-time.Hour))
path := getUserIdentityPath(expired.AccessKey, stsUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: expired, UpdatedAt: UTCNow().Add(-2 * time.Hour)}, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
var u UserIdentity
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("natural expiration retained a random key: %v", err)
}
// A retained early-revocation record is collectable only after the
// immutable token's expiration plus the skew allowance.
tomb := UserIdentity{Version: 1, Deleted: true, UpdatedAt: UTCNow().Add(-2 * time.Hour), ExpiresAt: expired.Expiration.Add(globalMaxSkewTime)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &tomb, path))
_ = sys.store.loadUser(ctx, expired.AccessKey, stsUser, make(map[string]UserIdentity))
if err := sys.store.loadIAMConfig(ctx, &u, path); !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS revocation not collected: %v", err)
}
if _, ok := sys.store.revisionIndex().snapshot()[path]; ok {
t.Fatal("expired STS retained an index entry")
}
})
}
}
func TestIAMPolicyDeletionRemainsExplicit(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
mustIAM(t, sys.DeletePolicy(ctx, "misspelled-policy", true))
r, err := loadIAMRevision(ctx, sys.store, getPolicyDocPath("misspelled-policy"))
mustIAM(t, err)
if r.Deleted {
t.Fatal("local nonexistent policy created a tombstone")
}
p, err := sys.store.GetPolicy("readwrite")
mustIAM(t, err)
if err := sys.DeletePolicy(ctx, "readwrite", true); err == nil {
t.Fatal("local pristine builtin policy became deletable")
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
mustIAM(t, sys.DeletePolicy(ctx, "readwrite", true))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if _, err := sys.store.GetPolicy("readwrite"); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("reload restored an explicitly deleted override: %v", err)
}
_, err = sys.SetPolicy(ctx, "readwrite", p)
mustIAM(t, err)
if _, err := sys.store.GetPolicy("readwrite"); err != nil {
t.Fatal("explicit policy recreation failed", err)
}
mustIAM(t, globalSiteReplicationSys.PeerAddPolicyHandler(ctx, "remote-unknown-policy", nil, UTCNow()))
r, err = loadIAMRevision(ctx, sys.store, getPolicyDocPath("remote-unknown-policy"))
mustIAM(t, err)
if !r.Deleted {
t.Fatal("replicated unknown deletion lost its version")
}
})
}
}
+65 -71
View File
@@ -26,7 +26,6 @@ import (
"sync"
"time"
jsoniter "github.com/json-iterator/go"
"github.com/minio/minio-go/v7/pkg/set"
"github.com/minio/minio/internal/config"
"github.com/minio/minio/internal/kms"
@@ -62,6 +61,7 @@ type IAMEtcdStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -69,13 +69,17 @@ type IAMEtcdStore struct {
}
func newIAMEtcdStore(client *etcd.Client, usersSysType UsersSysType) *IAMEtcdStore {
return &IAMEtcdStore{
store := &IAMEtcdStore{
iamCache: newIamCache(),
client: client,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (ies *IAMEtcdStore) revisionIndex() *iamRevisionIndex { return &ies.index }
func (ies *IAMEtcdStore) rlock() *iamCache {
ies.RLock()
return ies.iamCache
@@ -103,6 +107,7 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, itemPath),
@@ -111,24 +116,28 @@ func (ies *IAMEtcdStore) saveIAMConfig(ctx context.Context, item any, itemPath s
return err
}
}
return saveKeyEtcd(ctx, ies.client, itemPath, data, opts...)
if err := saveKeyEtcd(ctx, ies.client, itemPath, data, opts...); err != nil {
return err
}
ies.index.observe(itemPath, plain)
return nil
}
func getIAMConfig(item any, data []byte, itemPath string) error {
data, err := decryptData(data, itemPath)
func (ies *IAMEtcdStore) decodeIAMConfig(item any, data []byte, path string) error {
data, err := decryptData(data, path)
if err != nil {
return err
}
json := jsoniter.ConfigCompatibleWithStandardLibrary
ies.index.observe(path, data)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfig(ctx context.Context, item any, path string) error {
data, err := readKeyEtcd(ctx, ies.client, path)
data, err := ies.loadIAMConfigBytes(ctx, path)
if err != nil {
return err
}
return getIAMConfig(item, data, path)
return json.Unmarshal(data, item)
}
func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([]byte, error) {
@@ -136,11 +145,19 @@ func (ies *IAMEtcdStore) loadIAMConfigBytes(ctx context.Context, path string) ([
if err != nil {
return nil, err
}
return decryptData(data, path)
data, err = decryptData(data, path)
if err == nil {
ies.index.observe(path, data)
}
return data, err
}
func (ies *IAMEtcdStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteKeyEtcd(ctx, ies.client, path)
if err := deleteKeyEtcd(ctx, ies.client, path); err != nil {
return err
}
ies.index.forget(path)
return nil
}
func (ies *IAMEtcdStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, _ int) error {
@@ -162,6 +179,9 @@ func (ies *IAMEtcdStore) loadPolicyDoc(ctx context.Context, policy string, m map
return err
}
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -181,7 +201,11 @@ func (ies *IAMEtcdStore) getPolicyDocKV(ctx context.Context, kvs *mvccpb.KeyValu
return err
}
ies.index.observe(string(kvs.Key), data)
policy := extractPathPrefixAndSuffix(string(kvs.Key), iamConfigPoliciesPrefix, path.Base(string(kvs.Key)))
if p.Deleted {
return errNoSuchPolicy
}
m[policy] = p
return nil
}
@@ -207,7 +231,7 @@ func (ies *IAMEtcdStore) loadPolicyDocs(ctx context.Context, m map[string]Policy
func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue, userType IAMUserType, m map[string]UserIdentity, basePrefix string) error {
var u UserIdentity
err := getIAMConfig(&u, userkv.Value, string(userkv.Key))
err := ies.decodeIAMConfig(&u, userkv.Value, string(userkv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchUser
@@ -219,10 +243,14 @@ func (ies *IAMEtcdStore) getUserKV(ctx context.Context, userkv *mvccpb.KeyValue,
}
func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMUserType, u UserIdentity, m map[string]UserIdentity) error {
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
}
return errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, ies, getUserIdentityPath(user, userType))
return nil
}
if u.Credentials.AccessKey == "" {
@@ -231,16 +259,17 @@ func (ies *IAMEtcdStore) addUser(ctx context.Context, user string, userType IAMU
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
deleteKeyEtcd(ctx, ies.client, getUserIdentityPath(user, userType))
deleteKeyEtcd(ctx, ies.client, getMappedPolicyPath(user, userType, false))
}
// A temporarily unavailable signing key is not proof of expiration.
return nil
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, ies, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return errNoSuchUser
}
return err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
}
@@ -258,6 +287,9 @@ func (ies *IAMEtcdStore) loadSecretKey(ctx context.Context, user string, userTyp
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -274,6 +306,7 @@ func (ies *IAMEtcdStore) loadUser(ctx context.Context, user string, userType IAM
}
func (ies *IAMEtcdStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -312,6 +345,9 @@ func (ies *IAMEtcdStore) loadGroup(ctx context.Context, group string, m map[stri
}
return err
}
if gi.Deleted {
return errNoSuchGroup
}
m[group] = gi
return nil
}
@@ -349,13 +385,16 @@ func (ies *IAMEtcdStore) loadMappedPolicy(ctx context.Context, name string, user
}
return err
}
if !ies.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
func (ies *IAMEtcdStore) getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy], basePrefix string) error {
var p MappedPolicy
err := getIAMConfig(&p, kv.Value, string(kv.Key))
err := ies.decodeIAMConfig(&p, kv.Value, string(kv.Key))
if err != nil {
if err == errConfigNotFound {
return errNoSuchPolicy
@@ -363,6 +402,9 @@ func getMappedPolicy(kv *mvccpb.KeyValue, m *xsync.MapOf[string, MappedPolicy],
return err
}
name := extractPathPrefixAndSuffix(string(kv.Key), basePrefix, ".json")
if !ies.index.mappingAllowed(string(kv.Key), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -392,61 +434,13 @@ func (ies *IAMEtcdStore) loadMappedPolicies(ctx context.Context, userType IAMUse
// Parse all policies mapping to create the proper data model
for _, kv := range r.Kvs {
if err = getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
if err = ies.getMappedPolicy(kv, m, basePrefix); err != nil && !errors.Is(err, errNoSuchPolicy) {
return err
}
}
return nil
}
func (ies *IAMEtcdStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return ies.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (ies *IAMEtcdStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return ies.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (ies *IAMEtcdStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return ies.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (ies *IAMEtcdStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return ies.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (ies *IAMEtcdStore) deletePolicyDoc(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := ies.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (ies *IAMEtcdStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := ies.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (ies *IAMEtcdStore) deleteGroupInfo(ctx context.Context, name string) error {
err := ies.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
func (ies *IAMEtcdStore) watch(ctx context.Context, keyPath string) <-chan iamWatchEvent {
ch := make(chan iamWatchEvent)
+164
View File
@@ -0,0 +1,164 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"maps"
"slices"
"time"
"github.com/minio/minio-go/v7/pkg/set"
)
type (
iamGroupGrantsKey struct{}
iamGroupMutationKey struct{}
iamGroupMutation struct {
Members []string
Remove bool
StatusOnly bool
}
)
// Merge the intended mutation with the record read under the distributed
// revision lock, not the older cache used to prepare the request.
func mergeIAMGroupMutation(ctx context.Context, previous GroupInfo, next *GroupInfo) {
op, ok := ctx.Value(iamGroupMutationKey{}).(iamGroupMutation)
if !ok || previous.Deleted || previous.Version == 0 {
return
}
members := set.CreateStringSet(previous.Members...)
grants := maps.Clone(previous.MemberGrants)
if grants == nil {
grants = make(map[string]time.Time)
}
switch {
case op.StatusOnly:
// Only the status changes.
case op.Remove:
for _, member := range op.Members {
members.Remove(member)
delete(grants, member)
}
next.Status = previous.Status
default:
requested := set.CreateStringSet(next.Members...)
for _, member := range op.Members {
if !requested.Contains(member) {
continue
}
at := next.MemberGrants[member]
if at.Before(grants[member]) {
continue
}
members.Add(member)
grants[member] = at
}
next.Status = previous.Status
}
next.Members, next.MemberGrants = members.ToSlice(), grants
slices.Sort(next.Members)
}
// A non-nil map is supplied by the versioned peer envelope, including for
// snapshots. Missing times are unknown, never the snapshot's newer timestamp.
func withIAMGroupGrants(ctx context.Context, grants map[string]time.Time) context.Context {
return context.WithValue(ctx, iamGroupGrantsKey{}, grants)
}
func (c *iamCache) effectiveGroupMembers(gi GroupInfo) []string {
var members []string
for _, member := range gi.Members {
if c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
members = append(members, member)
}
}
return members
}
func (c *iamCache) effectiveUserGroups(user string) []string {
var groups []string
for group := range c.iamUserGroupMemberships[user] {
gi, ok := c.iamGroupsMap[group]
r := c.revisions.get(getGroupInfoPath(group))
if r.RevokedBefore.After(gi.RevokedBefore) {
gi.RevokedBefore = r.RevokedBefore
}
if ok && !r.Deleted && c.groupMemberAllowed(user, gi.MemberGrants[user], gi.RevokedBefore) {
groups = append(groups, group)
}
}
return groups
}
func (c *iamCache) addGroupMembers(ctx context.Context, gi GroupInfo, members []string) (GroupInfo, error) {
grants, versioned := ctx.Value(iamGroupGrantsKey{}).(map[string]time.Time)
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(gi.RevokedBefore) {
gi.RevokedBefore = boundary
}
origin, replicated := iamReplicationTime(ctx)
gi.Members = slices.Clone(gi.Members)
gi.MemberGrants = maps.Clone(gi.MemberGrants)
if gi.MemberGrants == nil {
gi.MemberGrants = make(map[string]time.Time)
}
current := set.CreateStringSet(gi.Members...)
gi.UpdatedAt = UTCNow()
if replicated {
gi.UpdatedAt = origin
}
for _, member := range members {
at := gi.UpdatedAt
r := c.userRevocation(member)
if replicated {
switch {
case versioned:
at = grants[member]
if at.After(origin) {
return gi, errInvalidArgument
}
case !r.RevokedBefore.IsZero() || r.Deleted || !gi.RevokedBefore.IsZero():
// Legacy snapshots cannot prove a post-revocation grant.
continue
case current.Contains(member):
continue
}
if !c.groupMemberAllowed(member, at, gi.RevokedBefore) {
continue
}
} else {
if current.Contains(member) && c.groupMemberAllowed(member, gi.MemberGrants[member], gi.RevokedBefore) {
continue // Editing the group is not reissuing every grant.
}
if !at.After(gi.RevokedBefore) {
at = gi.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(r.RevokedBefore) {
at = r.RevokedBefore.Add(time.Nanosecond)
}
if !at.After(gi.MemberGrants[member]) {
at = gi.MemberGrants[member].Add(time.Nanosecond)
}
if at.After(gi.UpdatedAt) {
gi.UpdatedAt = at
}
}
u, ok := c.iamUsersMap[member]
if !ok {
return gi, errNoSuchUser
}
if u.Credentials.IsTemp() || u.Credentials.IsServiceAccount() {
return gi, errIAMActionNotAllowed
}
if previous := gi.MemberGrants[member]; previous.After(at) {
continue
}
current.Add(member)
gi.MemberGrants[member] = at
}
gi.Members = current.ToSlice()
slices.Sort(gi.Members)
return gi, nil
}
+75
View File
@@ -0,0 +1,75 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"testing"
"time"
)
func TestIAMHealingResumesAfterLeadershipLoss(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
t.Cleanup(func() {
cancel()
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
first, loseFirst := context.WithCancel(ctx)
defer loseFirst()
select {
case locks <- LockContext{ctx: first}:
case <-time.After(time.Second):
t.Fatal("healer did not acquire its first leader context")
}
loseFirst() // A transient quorum loss cancels the distributed lease.
select {
case <-done:
t.Fatal("healer permanently exited after temporary leadership loss")
case locks <- LockContext{ctx: ctx}:
case <-time.After(time.Second):
t.Fatal("healer did not wait for reacquired leadership")
}
// Shutdown must also interrupt the wait for leadership after lease loss.
cancel()
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer did not stop with its owning context")
}
}
func TestIAMHealingLeadershipWaitCancels(t *testing.T) {
previous := globalLeaderLock
locks := make(chan LockContext)
globalLeaderLock = &sharedLock{lockContext: locks}
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
c := &SiteReplicationSys{}
go func() { c.startHealRoutine(ctx, nil); close(done) }()
cancel()
t.Cleanup(func() {
select {
case <-done:
case locks <- LockContext{ctx: ctx}:
}
<-done
globalLeaderLock = previous
})
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("healer ignored shutdown while waiting for leadership")
}
}
+60
View File
@@ -0,0 +1,60 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
// Measures steady-state index traversal, sorting and the capability request.
// The network peer acknowledges real batches but performs no disk I/O; this
// benchmark deliberately does not claim durable catch-up throughput.
func BenchmarkIAMRevisionConvergedHealing(b *testing.B) {
for _, n := range []int{1000, 10000} {
b.Run(fmt.Sprint(n), func(b *testing.B) {
ctx, sys, _ := prepareIAMRevisionFixture(b)
_, err := sys.CreateUser(ctx, "benchmark-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(b, err)
for i := range n {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
mustIAM(b, err)
sys.store.revisionIndex().observe(getUserIdentityPath(fmt.Sprintf("deleted-%06d", i), regUser), data)
}
var puts atomic.Int64
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
if r.Method == http.MethodPut {
puts.Add(1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: "benchmark-peer", Digest: "constant"}})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "benchmark-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
mustIAM(b, c.healIAMDeletions(ctx))
before := puts.Load()
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
mustIAM(b, c.healIAMDeletions(ctx))
}
b.StopTimer()
b.ReportMetric(float64(puts.Load()-before)/float64(b.N), "PUT/op")
if puts.Load() != before {
b.Fatal("steady-state healing replayed acknowledged records")
}
})
}
}
+56 -61
View File
@@ -45,6 +45,7 @@ type IAMObjectStore struct {
sync.RWMutex
*iamCache
index iamRevisionIndex
usersSysType UsersSysType
@@ -52,13 +53,17 @@ type IAMObjectStore struct {
}
func newIAMObjectStore(objAPI ObjectLayer, usersSysType UsersSysType) *IAMObjectStore {
return &IAMObjectStore{
store := &IAMObjectStore{
iamCache: newIamCache(),
objAPI: objAPI,
usersSysType: usersSysType,
}
store.revisions = &store.index
return store
}
func (iamOS *IAMObjectStore) revisionIndex() *iamRevisionIndex { return &iamOS.index }
func (iamOS *IAMObjectStore) rlock() *iamCache {
iamOS.RLock()
return iamOS.iamCache
@@ -87,6 +92,7 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
if err != nil {
return err
}
plain := data
if GlobalKMS != nil {
data, err = config.EncryptBytes(GlobalKMS, data, kms.Context{
minioMetaBucket: path.Join(minioMetaBucket, objPath),
@@ -95,7 +101,11 @@ func (iamOS *IAMObjectStore) saveIAMConfig(ctx context.Context, item any, objPat
return err
}
}
return saveConfig(ctx, iamOS.objAPI, objPath, data)
if err := saveConfig(ctx, iamOS.objAPI, objPath, data); err != nil {
return err
}
iamOS.index.observe(objPath, plain)
return nil
}
func decryptData(data []byte, objPath string) ([]byte, error) {
@@ -133,6 +143,7 @@ func (iamOS *IAMObjectStore) loadIAMConfigBytesWithMetadata(ctx context.Context,
if err != nil {
return nil, meta, err
}
iamOS.index.observe(objPath, data)
return data, meta, nil
}
@@ -146,7 +157,11 @@ func (iamOS *IAMObjectStore) loadIAMConfig(ctx context.Context, item any, objPat
}
func (iamOS *IAMObjectStore) deleteIAMConfig(ctx context.Context, path string) error {
return deleteConfig(ctx, iamOS.objAPI, path)
if err := deleteConfig(ctx, iamOS.objAPI, path); err != nil {
return err
}
iamOS.index.forget(path)
return nil
}
func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy string, m map[string]PolicyDoc, retries int) error {
@@ -171,6 +186,10 @@ func (iamOS *IAMObjectStore) loadPolicyDocWithRetry(ctx context.Context, policy
return err
}
if p.Deleted {
return errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -200,6 +219,10 @@ func (iamOS *IAMObjectStore) loadPolicy(ctx context.Context, policy string) (Pol
return p, err
}
if p.Deleted {
return PolicyDoc{}, errNoSuchPolicy
}
if p.Version == 0 {
// This means that policy was in the old version (without any
// timestamp info). We fetch the mod time of the file and save
@@ -245,6 +268,9 @@ func (iamOS *IAMObjectStore) loadSecretKey(ctx context.Context, user string, use
}
return "", err
}
if u.Deleted {
return "", errNoSuchUser
}
return u.Credentials.SecretKey, nil
}
@@ -258,10 +284,15 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
return u, err
}
if u.Deleted {
if userType == stsUser && !u.ExpiresAt.IsZero() && UTCNow().After(u.ExpiresAt) {
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
}
return UserIdentity{}, errNoSuchUser
}
if u.Credentials.IsExpired() {
// Delete expired identity - ignoring errors here.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
bestEffortIAMExpiration(ctx, iamOS, getUserIdentityPath(user, userType))
return u, errNoSuchUser
}
@@ -272,16 +303,18 @@ func (iamOS *IAMObjectStore) loadUserIdentity(ctx context.Context, user string,
if u.Credentials.SessionToken != "" {
jwtClaims, err := extractJWTClaims(u)
if err != nil {
if u.Credentials.IsTemp() {
// We should delete such that the client can re-request
// for the expiring credentials.
iamOS.deleteIAMConfig(ctx, getUserIdentityPath(user, userType))
iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(user, userType, false))
}
return u, errNoSuchUser
// During startup the site signing key may not be available yet.
// Reject this load without deleting a credential that has not expired.
return UserIdentity{}, errNoSuchUser
}
u.Credentials.Claims = jwtClaims.Map()
}
if err := checkIAMParentRevision(ctx, iamOS, u.Credentials); err != nil {
if errors.Is(err, errIAMStaleUpdate) {
return UserIdentity{}, errNoSuchUser
}
return UserIdentity{}, err
}
if u.Credentials.Description == "" {
u.Credentials.Description = u.Credentials.Comment
@@ -320,6 +353,7 @@ func (iamOS *IAMObjectStore) loadUser(ctx context.Context, user string, userType
}
func (iamOS *IAMObjectStore) loadUsers(ctx context.Context, userType IAMUserType, m map[string]UserIdentity) error {
ctx = withIAMExpirationCleanup(ctx)
var basePrefix string
switch userType {
case svcUser:
@@ -354,6 +388,9 @@ func (iamOS *IAMObjectStore) loadGroup(ctx context.Context, group string, m map[
}
return err
}
if g.Deleted {
return errNoSuchGroup
}
m[group] = g
return nil
}
@@ -391,6 +428,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyWithRetry(ctx context.Context, name
goto retry
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return errNoSuchPolicy
}
m.Store(name, p)
return nil
}
@@ -405,6 +445,9 @@ func (iamOS *IAMObjectStore) loadMappedPolicyInternal(ctx context.Context, name
}
return p, err
}
if !iamOS.index.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), p) {
return MappedPolicy{}, errNoSuchPolicy
}
return p, nil
}
@@ -824,54 +867,6 @@ func (iamOS *IAMObjectStore) loadAllFromObjStore(ctx context.Context, cache *iam
return nil
}
func (iamOS *IAMObjectStore) savePolicyDoc(ctx context.Context, policyName string, p PolicyDoc) error {
return iamOS.saveIAMConfig(ctx, &p, getPolicyDocPath(policyName))
}
func (iamOS *IAMObjectStore) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp MappedPolicy, opts ...options) error {
return iamOS.saveIAMConfig(ctx, mp, getMappedPolicyPath(name, userType, isGroup), opts...)
}
func (iamOS *IAMObjectStore) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u UserIdentity, opts ...options) error {
return iamOS.saveIAMConfig(ctx, u, getUserIdentityPath(name, userType), opts...)
}
func (iamOS *IAMObjectStore) saveGroupInfo(ctx context.Context, name string, gi GroupInfo) error {
return iamOS.saveIAMConfig(ctx, gi, getGroupInfoPath(name))
}
func (iamOS *IAMObjectStore) deletePolicyDoc(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getPolicyDocPath(name))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
err := iamOS.deleteIAMConfig(ctx, getMappedPolicyPath(name, userType, isGroup))
if err == errConfigNotFound {
err = errNoSuchPolicy
}
return err
}
func (iamOS *IAMObjectStore) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
err := iamOS.deleteIAMConfig(ctx, getUserIdentityPath(name, userType))
if err == errConfigNotFound {
err = errNoSuchUser
}
return err
}
func (iamOS *IAMObjectStore) deleteGroupInfo(ctx context.Context, name string) error {
err := iamOS.deleteIAMConfig(ctx, getGroupInfoPath(name))
if err == errConfigNotFound {
err = errNoSuchGroup
}
return err
}
// Lists objects in the minioMetaBucket at the given path prefix. All returned
// items have the pathPrefix removed from their names.
func listIAMConfigItems(ctx context.Context, objAPI ObjectLayer, pathPrefix string) <-chan itemOrErr[string] {
+131
View File
@@ -0,0 +1,131 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"fmt"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// Uses APIs shared with the pre-revision tree so the same benchmark can be
// overlaid on that tree for a comparable local baseline.
func prepareIAMPerformanceFixture(b *testing.B) (context.Context, *IAMSys) {
b.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
if err != nil {
b.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
b.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.initStore(obj, nil)
if err := globalIAMSys.Load(ctx, true); err != nil {
b.Fatal(err)
}
b.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys
}
func BenchmarkIAMCachedCredential(b *testing.B) {
for _, kind := range []string{"user", "service", "sts"} {
b.Run(kind, func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
key := parent
if kind == "service" {
c, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "benchmark-service", secretKey: "benchmark-service-password"})
if err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
if kind == "sts" {
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
c, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
c.ParentUser = parent
if _, err := sys.SetTempUser(ctx, c.AccessKey, c, ""); err != nil {
b.Fatal(err)
}
key = c.AccessKey
}
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, ok := sys.store.GetUser(key); !ok {
b.Fatal("credential missing")
}
}
})
}
}
func BenchmarkIAMSetTempUser(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
const parent = "benchmark-sts-parent"
_, err := sys.CreateUser(ctx, parent, madmin.AddOrUpdateUserReq{SecretKey: "benchmark-user-password", Status: madmin.AccountEnabled})
if err != nil {
b.Fatal(err)
}
secret, err := getTokenSigningKey()
if err != nil {
b.Fatal(err)
}
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
if err != nil {
b.Fatal(err)
}
cred.ParentUser = parent
b.ReportAllocs()
b.ResetTimer()
for b.Loop() {
if _, err := sys.SetTempUser(ctx, cred.AccessKey, cred, "readwrite"); err != nil {
b.Fatal(err)
}
}
}
// Run with -benchtime=1x. Preparation is outside the timer; each measured load
// sees a fresh set of expired reusable service-account records.
func BenchmarkIAMColdLoadExpiredServices(b *testing.B) {
for _, count := range []int{100, 1000} {
b.Run(fmt.Sprint(count), func(b *testing.B) {
ctx, sys := prepareIAMPerformanceFixture(b)
b.ReportAllocs()
for i := 0; i < b.N; i++ {
b.StopTimer()
for j := 0; j < count; j++ {
key := fmt.Sprintf("expired-benchmark-%d-%d", i, j)
u := UserIdentity{Version: 1, UpdatedAt: UTCNow().Add(-2 * time.Hour), Credentials: auth.Credentials{AccessKey: key, SecretKey: "expired-benchmark-password", ParentUser: "absent-idp-parent", Expiration: UTCNow().Add(-time.Hour), Status: auth.AccountOn}}
if err := sys.store.saveIAMConfig(ctx, &u, getUserIdentityPath(key, svcUser)); err != nil {
b.Fatal(err)
}
}
b.StartTimer()
if err := sys.store.LoadIAMCache(ctx, true); err != nil {
b.Fatal(err)
}
}
})
}
}
+168
View File
@@ -0,0 +1,168 @@
package cmd
import (
"context"
"errors"
"os"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/pgsty/silo-pkg/v3/policy"
)
func TestReviewIAMRevokedUserReplay(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
policyAt, err := globalIAMSys.PolicyDBSet(ctx, user, "readwrite", regUser, false)
if err != nil {
t.Fatal(err)
}
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "review-bucket", ObjectName: "review-object"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("seed must allow object read")
}
if err := globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("deletion did not remove initial permission")
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Errorf("revoked user restored by an older replicated create, GetUserInfo error = %v", err)
}
if err := globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: user, UserType: int(regUser), Policy: "readwrite"}, policyAt); err != nil {
t.Fatal(err)
}
if globalIAMSys.IsAllowed(args) {
t.Error("older replicated identity and policy events restored revoked S3 read permission")
}
}
func TestReviewIAMSourceTimestampOrder(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "review-ordered-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "review-valid-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour)
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin); err != nil {
t.Fatal(err)
}
if err := globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(time.Minute)); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("newer source deletion skipped after delayed creation, GetUserInfo error = %v", err)
}
}
// A user's old group grant must not return after deletion and deliberate recreation.
func TestR3CandidateOldGroupReplayAfterRecreation(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-group-member", "r3-granting-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("recreation must start without deleted group membership")
}
must(peer.PeerGroupInfoChangeHandler(ctx, add, origin.Add(time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if globalIAMSys.IsAllowed(args) {
t.Fatal("old group event restored the deleted user's read grant after recreation and durable reload")
}
}
// The user delete is also a revocation of its earlier group memberships.
func TestR3CandidateLateDeleteRetainsOldGroupGrant(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user, group := "r3-late-member", "r3-late-group"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-r3-user-password", Status: madmin.AccountEnabled}
peer := &globalSiteReplicationSys
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin))
must(peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin.Add(time.Minute)))
must(peer.PeerPolicyMappingHandler(ctx, &madmin.SRPolicyMapping{UserOrGroup: group, IsGroup: true, UserType: int(regUser), Policy: "readwrite"}, origin.Add(time.Minute)))
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "r3-bucket", ObjectName: "probe"}
if !globalIAMSys.IsAllowed(args) {
t.Fatal("fixture must grant through group")
}
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, origin.Add(3*time.Minute)))
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(globalIAMSys.store.LoadIAMCache(ctx, false))
if _, ok := globalIAMSys.GetUser(ctx, user); !ok {
t.Fatal("newer identity must survive")
}
if globalIAMSys.IsAllowed(args) {
t.Fatal("late user deletion retained the older group grant on the recreated identity")
}
}
+321
View File
@@ -0,0 +1,321 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"sort"
"strings"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
xhttp "github.com/minio/minio/internal/http"
"github.com/pgsty/silo-pkg/v3/policy"
)
const (
iamRevisionProtocol = 1
iamRevisionPeerPath = "/v3/site-replication/peer/iam-revisions"
iamUserBoundaryType = "silo-user-revocation"
iamGroupBoundaryType = "silo-group-revocation"
maxIAMRevisionBatch = 128
)
var iamRevisionInstance = mustGetUUID()
type iamUserBoundary struct {
User string `json:"user"`
Before time.Time `json:"before"`
}
type iamGroupBoundary struct {
Group string `json:"group"`
Before time.Time `json:"before"`
}
// The server owns this additive protocol, without changing the client SDK or
// overloading a policy/document field. Old servers reject the dedicated route
// before applying any change that would lose revocation or member metadata.
type iamReplicationItem struct {
madmin.SRIAMItem
GroupGrants map[string]time.Time `json:"groupGrants,omitempty"`
GroupSnapshot bool `json:"groupSnapshot,omitempty"`
UserRevocation *iamUserBoundary `json:"userRevocation,omitempty"`
GroupRevocation *iamGroupBoundary `json:"groupRevocation,omitempty"`
RevokedBefore time.Time `json:"revokedBefore,omitempty"`
}
type iamRevisionBatch struct {
Version int `json:"version"`
Items []iamReplicationItem `json:"items"`
}
type iamRevisionStatus struct {
Version int `json:"version"`
Node string `json:"node"`
Instance string `json:"instance"`
Digest string `json:"digest"`
}
type iamRevisionResponse struct {
iamRevisionStatus
Errors []string `json:"errors,omitempty"`
}
type iamRevisionBatchError struct{ failures []string }
func (e *iamRevisionBatchError) Error() string {
return "IAM revision batch: " + strings.Join(e.failures, "; ")
}
type iamRevisionProgress struct {
Instances map[string]string
Acknowledged map[string]string
}
type iamRevisionMetrics struct {
healFailures atomic.Uint64
healLastSuccess atomic.Int64
healDurationMillis atomic.Int64
}
func iamRevisionDigest(items map[string]iamRevision) string {
paths := make([]string, 0, len(items))
for path := range items {
paths = append(paths, path)
}
sort.Strings(paths)
h := sha256.New()
for _, path := range paths {
r := items[path]
fmt.Fprintf(h, "%q %s %t %s\n", path, r.timestamp().UTC().Format(time.RFC3339Nano), r.Deleted, r.RevokedBefore.UTC().Format(time.RFC3339Nano))
}
return hex.EncodeToString(h.Sum(nil))
}
func (store *IAMStoreSys) iamRevisionStatus() iamRevisionStatus {
node := globalLocalNodeName
if node == "" {
node = "local"
}
return iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: iamRevisionInstance, Digest: store.revisionIndex().digest()}
}
func executeIAMRevisionRequest(ctx context.Context, client *madmin.AdminClient, method string, batch *iamRevisionBatch) (status iamRevisionStatus, err error) {
var content []byte
if batch != nil {
content, err = json.Marshal(batch)
if err != nil {
return status, err
}
}
resp, err := client.ExecuteMethod(ctx, method, madmin.RequestData{RelPath: iamRevisionPeerPath, QueryValues: url.Values{"api-version": {madmin.SiteReplAPIVersion}}, Content: content})
if resp != nil {
defer xhttp.DrainBody(resp.Body)
}
if err != nil {
return status, err
}
if resp.StatusCode != http.StatusOK {
var remote madmin.ErrorResponse
if json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&remote) == nil && remote.Code != "" {
return status, remote
}
return status, fmt.Errorf("IAM revision protocol requires upgraded peers: %s", resp.Status)
}
var response iamRevisionResponse
if err = json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(&response); err != nil {
return status, err
}
status = response.iamRevisionStatus
if status.Version != iamRevisionProtocol || status.Node == "" || status.Instance == "" || status.Digest == "" {
return status, errors.New("peer did not acknowledge the IAM revision protocol")
}
if len(response.Errors) != 0 {
return status, &iamRevisionBatchError{failures: response.Errors}
}
return status, nil
}
type (
iamRecordBoundaryKey struct{}
iamGroupSnapshotKey struct{}
)
func (c *SiteReplicationSys) replicationItem(ctx context.Context, item madmin.SRIAMItem) (iamReplicationItem, error) {
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemSvcAcc && item.SvcAccChange != nil {
var key string
if item.SvcAccChange.Create != nil {
key = item.SvcAccChange.Create.AccessKey
} else if item.SvcAccChange.Update != nil {
key = item.SvcAccChange.Update.AccessKey
}
if key != "" {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(key, svcUser))
if err != nil {
return out, err
}
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && !item.GroupInfo.UpdateReq.IsRemove {
out.GroupSnapshot = true
var gi GroupInfo
if err := globalIAMSys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath(item.GroupInfo.UpdateReq.Group)); err != nil {
return out, err
}
// The matching persisted snapshot carries member grant times. If a
// later write won before sending, propagate that whole newer state.
if gi.Deleted {
return out, errIAMStaleUpdate
}
out.UpdatedAt = gi.UpdatedAt
out.RevokedBefore = gi.RevokedBefore
out.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: item.GroupInfo.UpdateReq.Group, Status: madmin.GroupStatus(gi.Status)}}
cache := globalIAMSys.store.rlock()
out.GroupInfo.UpdateReq.Members = cache.effectiveGroupMembers(gi)
globalIAMSys.store.runlock()
out.GroupGrants = make(map[string]time.Time, len(out.GroupInfo.UpdateReq.Members))
for _, member := range out.GroupInfo.UpdateReq.Members {
out.GroupGrants[member] = gi.MemberGrants[member]
}
}
if item.Type == madmin.SRIAMItemGroupInfo && item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getGroupInfoPath(item.GroupInfo.UpdateReq.Group))
if err != nil {
return out, err
}
if !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
}
}
if item.Type == madmin.SRIAMItemIAMUser && item.IAMUser != nil {
r, err := loadIAMRevision(ctx, globalIAMSys.store, getUserIdentityPath(item.IAMUser.AccessKey, regUser))
if err != nil {
return out, err
}
if item.IAMUser.IsDeleteReq && !r.Deleted && !r.RevokedBefore.IsZero() {
out.Type = iamUserBoundaryType
out.IAMUser = nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
out.UpdatedAt = r.RevokedBefore
} else if !item.IAMUser.IsDeleteReq {
if r.Deleted || r.timestamp().After(item.UpdatedAt) {
return out, errIAMStaleUpdate
}
out.RevokedBefore = r.RevokedBefore
}
}
return out, nil
}
func applyIAMReplicationItem(ctx context.Context, item iamReplicationItem) error {
if item.GroupInfo != nil {
if item.GroupSnapshot {
ctx = context.WithValue(ctx, iamGroupSnapshotKey{}, true)
}
// A nil map also explicitly denotes unknown legacy grants. Do not
// turn an unrelated group edit into a new grant after a revocation.
ctx = withIAMGroupGrants(ctx, item.GroupGrants)
}
if !item.RevokedBefore.IsZero() {
if item.RevokedBefore.After(item.UpdatedAt) {
return errSRInvalidRequest(errInvalidArgument)
}
ctx = context.WithValue(ctx, iamRecordBoundaryKey{}, item.RevokedBefore)
}
switch item.Type {
case iamUserBoundaryType:
if item.UserRevocation == nil || item.UserRevocation.User == "" || item.UserRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
return iamReplicationError(globalIAMSys.DeleteUser(withIAMReplicationTime(ctx, item.UserRevocation.Before), item.UserRevocation.User, true))
case iamGroupBoundaryType:
if item.GroupRevocation == nil || item.GroupRevocation.Group == "" || item.GroupRevocation.Before.IsZero() {
return errSRInvalidRequest(errInvalidArgument)
}
_, err := globalIAMSys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, item.GroupRevocation.Before), item.GroupRevocation.Group, nil)
return iamReplicationError(err)
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, nil, item.UpdatedAt)
}
p, err := policy.ParseConfig(bytes.NewReader(item.Policy))
if err != nil {
return err
}
if p.IsEmpty() {
p = nil
}
return globalSiteReplicationSys.PeerAddPolicyHandler(ctx, item.Name, p, item.UpdatedAt)
case madmin.SRIAMItemSvcAcc:
return globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, item.SvcAccChange, item.UpdatedAt)
case madmin.SRIAMItemPolicyMapping:
return globalSiteReplicationSys.PeerPolicyMappingHandler(ctx, item.PolicyMapping, item.UpdatedAt)
case madmin.SRIAMItemSTSAcc:
return globalSiteReplicationSys.PeerSTSAccHandler(ctx, item.STSCredential, item.UpdatedAt)
case madmin.SRIAMItemIAMUser:
return globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt)
case madmin.SRIAMItemGroupInfo:
return globalSiteReplicationSys.PeerGroupInfoChangeHandler(ctx, item.GroupInfo, item.UpdatedAt)
default:
return errSRInvalidRequest(errInvalidArgument)
}
}
func (a adminAPIHandlers) SRPeerIAMRevisions(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
if obj, _ := validateAdminReq(ctx, w, r, policy.SiteReplicationOperationAction); obj == nil {
return
}
var failures []string
if r.Method == http.MethodPut {
var batch iamRevisionBatch
if err := parseJSONBody(ctx, r.Body, &batch, ""); err != nil {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, err), r.URL)
return
}
if batch.Version != iamRevisionProtocol || len(batch.Items) == 0 || len(batch.Items) > maxIAMRevisionBatch {
writeErrorResponseJSON(ctx, w, toAdminAPIErr(ctx, errSRInvalidRequest(errInvalidArgument)), r.URL)
return
}
for i, item := range batch.Items {
if err := applyIAMReplicationItem(ctx, item); err != nil {
failures = append(failures, fmt.Sprintf("item %d (%s): %v", i, item.Type, err))
}
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: globalIAMSys.store.iamRevisionStatus(), Errors: failures})
}
// A site endpoint can balance requests across nodes sharing durable IAM state.
// Switching between known node incarnations preserves ACKs; a new incarnation
// conservatively invalidates them so restoring an old backend cannot inherit
// acknowledgements from before the restore.
func (p *iamRevisionProgress) observePeer(status iamRevisionStatus) {
if p.Instances == nil {
p.Instances = make(map[string]string)
}
if p.Instances[status.Node] != status.Instance || p.Acknowledged == nil {
p.Instances[status.Node] = status.Instance
p.Acknowledged = make(map[string]string)
}
}
+161
View File
@@ -0,0 +1,161 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
)
func TestIAMRevisionProtocolDoesNotFallBackToLegacy(t *testing.T) {
var requests atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requests.Add(1)
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unsafe fallback path: %s", r.URL.Path)
}
w.WriteHeader(http.StatusNotFound)
_, _ = w.Write([]byte(`{"Code":"NotImplemented","Message":"old server"}`))
}))
defer server.Close()
client, err := madmin.New(strings.TrimPrefix(server.URL, "http://"), "test-access", "valid-test-secret", false)
mustIAM(t, err)
_, err = executeIAMRevisionRequest(context.Background(), client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: []iamReplicationItem{{SRIAMItem: madmin.SRIAMItem{Type: iamUserBoundaryType}, UserRevocation: &iamUserBoundary{User: "recreated", Before: UTCNow()}}}})
if err == nil || requests.Load() != 1 {
t.Fatalf("old peer must reject without fallback, err=%v requests=%d", err, requests.Load())
}
}
type iamNoHealingScanStore struct{ IAMStorageAPI }
func (s *iamNoHealingScanStore) listIAMConfigPaths(context.Context) ([]string, error) {
panic("healing must use the loaded revision index")
}
func TestIAMRevisionHealingAcknowledgements(t *testing.T) {
for _, balanced := range []bool{false, true} {
t.Run(fmt.Sprintf("load_balanced_%t", balanced), func(t *testing.T) { testIAMRevisionHealingAcknowledgements(t, balanced) })
}
}
func testIAMRevisionHealingAcknowledgements(t *testing.T, balanced bool) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "ack-sync", madmin.AddOrUpdateUserReq{SecretKey: "valid-sync-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
for i := range maxIAMRevisionBatch*2 + 1 {
at := UTCNow().Add(time.Duration(i) * time.Nanosecond)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath(fmt.Sprintf("ack-%04d", i), regUser)))
}
sys.store.IAMStorageAPI = &iamNoHealingScanStore{IAMStorageAPI: sys.store.IAMStorageAPI}
var mu sync.Mutex
var puts, gets int
var applied int
instance := "boot-1"
failSecondBatch := true
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
w.WriteHeader(http.StatusOK)
return
}
mu.Lock()
defer mu.Unlock()
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" {
t.Errorf("unexpected request: %s", r.URL.Path)
w.WriteHeader(404)
return
}
var failures []string
if r.Method == http.MethodGet {
gets++
} else {
puts++
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(400)
return
}
if len(batch.Items) > maxIAMRevisionBatch {
t.Error("batch exceeds limit")
}
if failSecondBatch && puts == 2 {
failures = []string{"injected item error"}
} else {
applied += len(batch.Items)
}
}
node := "node-1"
if balanced {
node = fmt.Sprintf("node-%d", (gets+puts)%2+1)
}
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: node, Instance: node + instance, Digest: fmt.Sprintf("%d", applied)}, Errors: failures})
}))
defer server.Close()
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "ack-sync", Peers: map[string]madmin.PeerInfo{globalDeploymentID(): {DeploymentID: globalDeploymentID(), Name: "local"}, "remote": {DeploymentID: "remote", Name: "remote", Endpoint: server.URL}}}}
if err := c.healIAMDeletions(ctx); err == nil {
t.Fatal("item failure was hidden")
}
mu.Lock()
if puts != 3 || applied != maxIAMRevisionBatch+1 {
t.Errorf("failed middle batch blocked later revocations: puts=%d applied=%d", puts, applied)
}
failSecondBatch = false
applied++ // Unrelated remote mutation changes its digest.
mu.Unlock()
at := UTCNow()
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Deleted: true, UpdatedAt: at, RevokedBefore: at}, getUserIdentityPath("ack-new-local", regUser)))
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 {
t.Errorf("did not resume at unacknowledged batch: puts=%d", puts)
}
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
if puts != 5 || gets != 3 {
t.Errorf("converged records were replayed: puts=%d gets=%d", puts, gets)
}
instance = "boot-2"
mu.Unlock()
mustIAM(t, c.healIAMDeletions(ctx))
mu.Lock()
defer mu.Unlock()
if puts != 8 {
t.Fatalf("peer restart reused an old acknowledgement: puts=%d", puts)
}
}
func TestIAMRevisionIndexRebuildsFromStorage(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "index-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-parent-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(ctx, user, req)
mustIAM(t, err)
mustIAM(t, sys.DeleteUser(ctx, user, false))
before := sys.store.revisionIndex().snapshot()
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, store.LoadIAMCache(ctx, true))
if iamRevisionDigest(before) != iamRevisionDigest(store.revisionIndex().snapshot()) {
t.Fatal("ordinary IAM loading did not restore the deletion index")
}
_, err = store.AddUser(ctx, user, req)
mustIAM(t, err)
r := store.revisionIndex().get(getUserIdentityPath(user, regUser))
if r.Deleted || r.RevokedBefore.IsZero() {
t.Fatal("recreation discarded the retained boundary")
}
if r.Credentials.SecretKey != "" || r.Credentials.SessionToken != "" {
t.Fatal("index retained credentials")
}
}
+393
View File
@@ -0,0 +1,393 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"errors"
"fmt"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
"github.com/pgsty/silo-pkg/v3/policy"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
func prepareIAMRevisionFixture(t testing.TB, backend ...string) (context.Context, *IAMSys, ObjectLayer) {
t.Helper()
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
disks, err := getRandomDisks(1)
mustIAM(t, err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
mustIAM(t, err)
initAllSubsystems(ctx)
// Deliberately omit the periodic refresh goroutine. Fault injection can
// replace this fixture's storage interface without racing initialization.
var client *etcd.Client
if len(backend) != 0 && backend[0] == "etcd" {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
cancel()
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
client, err = etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
mustIAM(t, err)
prefix := fmt.Sprintf("/silo-boundary-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
t.Cleanup(func() { client.Delete(context.Background(), "", etcd.WithPrefix()); client.Close() })
}
globalIAMSys.initStore(obj, client)
mustIAM(t, globalIAMSys.Load(ctx, true))
t.Cleanup(func() { cancel(); obj.Shutdown(context.Background()); os.RemoveAll(disks[0]); resetTestGlobals() })
return ctx, globalIAMSys, obj
}
func mustIAM(t testing.TB, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
var errIAMInjectedWrite = errors.New("injected IAM persistence failure")
type iamFailingCleanupStore struct {
IAMStorageAPI
parentPath string
beforeCommit bool
}
func (s *iamFailingCleanupStore) saveIAMConfig(ctx context.Context, item any, path string, opts ...options) error {
if s.beforeCommit || path != s.parentPath {
return errIAMInjectedWrite
}
return s.IAMStorageAPI.saveIAMConfig(ctx, item, path, opts...)
}
func TestIAMRevocationCommitBoundary(t *testing.T) {
for _, before := range []bool{true, false} {
name := "after_identity_commit"
if before {
name = "before_identity_commit"
}
t.Run(name, func(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
const user = "commit-boundary-user"
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(time.Minute)), user, "readwrite", regUser, false)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(time.Minute)), "commit-group", []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "commit-group", "readwrite", regUser, true)
mustIAM(t, err)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), user, nil, newServiceAccountOpts{accessKey: "commit-child", secretKey: "valid-child-password"})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no grant")
}
siblingStore := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, siblingStore.LoadIAMCache(ctx, true))
sibling := &IAMSys{store: siblingStore, usersSysType: MinIOUsersSysType}
tg, err := grid.SetupTestGrid(2)
mustIAM(t, err)
defer tg.Cleanup()
var notifications atomic.Int32
mustIAM(t, deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
notifications.Add(1)
if err := sibling.LoadUserAfterDelete(ctx, r.Get(peerRESTUser)); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
mustIAM(t, err)
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{host: host, gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) }}}}
original := sys.store.IAMStorageAPI
sys.store.IAMStorageAPI = &iamFailingCleanupStore{IAMStorageAPI: original, parentPath: getUserIdentityPath(user, regUser), beforeCommit: before}
boundary := origin.Add(2 * time.Minute)
err = sys.DeleteUser(withIAMReplicationTime(ctx, boundary), user, true)
if !errors.Is(err, errIAMInjectedWrite) {
t.Fatalf("expected write failure, got %v", err)
}
sys.store.IAMStorageAPI = original
r, err := loadIAMRevision(ctx, original, getUserIdentityPath(user, regUser))
mustIAM(t, err)
if before {
if r.Deleted || !sys.IsAllowed(args) || !sibling.IsAllowed(args) || notifications.Load() != 0 {
t.Fatal("failure before commit changed the identity or grant")
}
return
}
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("cleanup failure lost durable revocation")
}
if sys.IsAllowed(args) || sibling.IsAllowed(args) || notifications.Load() != 1 {
t.Fatal("cleanup failure retained old permission")
}
// Subsequent fixture writes need no additional RPC handlers.
globalNotificationSys = &NotificationSys{}
// Recreate after the partial cleanup. The old mapping, group member
// and child still exist in storage; none may authorize this identity.
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), user, req)
mustIAM(t, err)
reloaded := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
mustIAM(t, reloaded.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: reloaded, usersSysType: MinIOUsersSysType}
if fresh.IsAllowed(args) {
t.Fatal("cold reload restored partially cleaned-up grants")
}
if _, ok := reloaded.GetUser(child.AccessKey); ok {
t.Fatal("cold reload restored the old child")
}
gd, err := reloaded.GetGroupDescription("commit-group")
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("listing exposed a revoked group relation: %v", gd.Members)
}
_, err = sys.AddUsersToGroup(ctx, "commit-group", []string{user})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit new group grant was not accepted")
}
})
}
}
func TestIAMGroupGrantVersionsSurviveSnapshotsAndRecreation(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
for _, user := range []string{"grant-alice", "grant-bob"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, req)
mustIAM(t, err)
}
grant := origin.Add(time.Minute)
_, err := sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), "grant-group", []string{"grant-alice"})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "grant-group", "readwrite", regUser, true)
mustIAM(t, err)
boundary := origin.Add(2 * time.Minute)
mustIAM(t, sys.DeleteUser(withIAMReplicationTime(ctx, boundary), "grant-alice", false))
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), "grant-alice", req)
mustIAM(t, err)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), "grant-group", []string{"grant-bob"})
mustIAM(t, err)
_, err = sys.SetGroupStatus(withIAMReplicationTime(ctx, origin.Add(5*time.Minute)), "grant-group", true)
mustIAM(t, err)
var gi GroupInfo
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(grant) {
t.Fatal("unrelated group edits refreshed an old grant")
}
args := policy.Args{AccountName: "grant-alice", Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
for _, stale := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(6 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: "grant-group", Members: []string{"grant-alice", "grant-bob"}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{"grant-alice": stale, "grant-bob": origin.Add(4 * time.Minute)}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
if sys.IsAllowed(args) {
t.Fatalf("snapshot restored revoked grant %s", stale)
}
gd, err := sys.GetGroupDescription("grant-group")
mustIAM(t, err)
if len(gd.Members) != 1 || gd.Members[0] != "grant-bob" {
t.Fatalf("inconsistent effective members: %v", gd.Members)
}
}
// Only an explicit post-revocation grant restores access.
freshAt, err := sys.AddUsersToGroup(ctx, "grant-group", []string{"grant-alice"})
mustIAM(t, err)
if !sys.IsAllowed(args) {
t.Fatal("explicit regrant rejected")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, sys.store.loadIAMConfig(ctx, &gi, getGroupInfoPath("grant-group")))
if !gi.MemberGrants["grant-alice"].Equal(freshAt) {
t.Fatal("new grant version was not persisted")
}
if !gi.MemberGrants["grant-bob"].Equal(origin.Add(4 * time.Minute)) {
t.Fatal("regranting Alice changed Bob's grant")
}
}
func TestIAMGroupRevocationCommitAndRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) { testIAMGroupRevocationCommitAndRecreation(t, backend) })
}
}
func testIAMGroupRevocationCommitAndRecreation(t *testing.T, backend string) {
ctx, sys, obj := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
user, group := "group-boundary-user", "group-boundary"
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant, boundary := origin.Add(time.Minute), origin.Add(2*time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
// A newer mapping must not veto the authoritative group deletion.
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin.Add(3*time.Minute)), group, "readwrite", regUser, true)
mustIAM(t, err)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, boundary), group, nil)
mustIAM(t, err)
r, err := loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if !r.Deleted || !r.RevokedBefore.Equal(boundary) {
t.Fatal("newer mapping swallowed group deletion")
}
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, origin.Add(4*time.Minute)), group, nil)
mustIAM(t, err)
for _, at := range []time.Time{grant, boundary, {}} {
item := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, UpdatedAt: origin.Add(5 * time.Minute), GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}}, GroupSnapshot: true, GroupGrants: map[string]time.Time{user: at}}
mustIAM(t, applyIAMReplicationItem(ctx, item))
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if len(gd.Members) != 0 {
t.Fatalf("group recreation restored grant %s", at)
}
}
_, err = sys.AddUsersToGroup(ctx, group, []string{user})
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("explicit group regrant was rejected")
}
// The newer live snapshot may arrive before an older group deletion.
lateBoundary := origin.Add(6 * time.Minute)
_, err = sys.RemoveUsersFromGroup(withIAMReplicationTime(ctx, lateBoundary), group, nil)
mustIAM(t, err)
r, err = loadIAMRevision(ctx, sys.store, getGroupInfoPath(group))
mustIAM(t, err)
if r.Deleted || !r.RevokedBefore.Equal(lateBoundary) {
t.Fatal("late deletion lost the live group's revocation boundary")
}
// The old mapping is now revoked; a new explicit mapping restores access.
if sys.IsAllowed(args) {
t.Fatal("late group boundary retained an old mapping")
}
_, err = sys.PolicyDBSet(ctx, group, "readwrite", regUser, true)
mustIAM(t, err)
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
if es, ok := sys.store.IAMStorageAPI.(*IAMEtcdStore); ok {
store.IAMStorageAPI = newIAMEtcdStore(es.client, MinIOUsersSysType)
}
mustIAM(t, store.LoadIAMCache(ctx, true))
fresh := &IAMSys{store: store, usersSysType: MinIOUsersSysType}
if !fresh.IsAllowed(args) {
t.Fatal("reload lost explicit grants after a retained group boundary")
}
item, err := globalSiteReplicationSys.replicationItem(ctx, madmin.SRIAMItem{Type: madmin.SRIAMItemGroupInfo, GroupInfo: &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group}}, UpdatedAt: r.timestamp()})
mustIAM(t, err)
if !item.RevokedBefore.Equal(lateBoundary) || !item.GroupGrants[user].After(lateBoundary) {
t.Fatal("group snapshot lost revision metadata")
}
}
// A committed revision is observable before all cached dependents have been
// cleaned up. Every authorization read must apply that boundary in this window.
func TestIAMCachedMappingHonorsCommittedRevision(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
parent := "cached-external-parent"
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), parent, "readwrite", stsUser, false)
mustIAM(t, err)
policies, err := sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) == 0 {
t.Fatal("fixture has no STS-parent mapping")
}
mustIAM(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Deleted: true, UpdatedAt: origin.Add(time.Minute)}, getMappedPolicyPath(parent, stsUser, false)))
policies, err = sys.PolicyDBGet(parent)
mustIAM(t, err)
if len(policies) != 0 {
t.Fatal("cached STS mapping ignored its own namespace tombstone")
}
user, group := "cached-group-user", "cached-group"
_, err = sys.CreateUser(withIAMReplicationTime(ctx, origin), user, madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
grant := origin.Add(5 * time.Minute)
_, err = sys.AddUsersToGroup(withIAMReplicationTime(ctx, grant), group, []string{user})
mustIAM(t, err)
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), group, "readwrite", regUser, true)
mustIAM(t, err)
args := policy.Args{AccountName: user, Action: policy.GetObjectAction, BucketName: "bucket", ObjectName: "object"}
if !sys.IsAllowed(args) {
t.Fatal("fixture has no group grant")
}
// A late deletion preserves the newer member grant but revokes the older
// policy mapping. Simulate the interval before mapping cleanup completes.
gi := GroupInfo{Version: 1, Status: statusEnabled, Members: []string{user}, MemberGrants: map[string]time.Time{user: grant}, UpdatedAt: grant, RevokedBefore: origin.Add(2 * time.Minute)}
mustIAM(t, sys.store.saveIAMConfig(ctx, &gi, getGroupInfoPath(group)))
if sys.IsAllowed(args) {
t.Fatal("cached group mapping ignored the committed group boundary")
}
gd, err := sys.GetGroupDescription(group)
mustIAM(t, err)
if gd.Policy != "" {
t.Fatal("group listing exposed a revoked mapping")
}
}
type (
iamExpiryLockFailure struct {
ObjectLayer
path string
}
iamFailedExpiryLock struct{ RWLocker }
)
func (o *iamExpiryLockFailure) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path+".revision-lock" {
return &iamFailedExpiryLock{RWLocker: lock}
}
return lock
}
func (l *iamFailedExpiryLock) GetLock(context.Context, *dynamicTimeout) (LockContext, error) {
return LockContext{}, errIAMInjectedWrite
}
func TestIAMExpiredCredentialCleanupDoesNotBlockLoading(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "healthy-user", madmin.AddOrUpdateUserReq{SecretKey: "healthy-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, err = sys.PolicyDBSet(ctx, "healthy-user", "readwrite", regUser, false)
mustIAM(t, err)
c, _, err := sys.NewServiceAccount(ctx, "healthy-user", nil, newServiceAccountOpts{accessKey: "expired-service", secretKey: "expired-service-password"})
mustIAM(t, err)
c.Expiration = UTCNow().Add(-time.Hour)
path := getUserIdentityPath(c.AccessKey, svcUser)
mustIAM(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: c, UpdatedAt: UTCNow()}, path))
// A cold loader sees the existing version but cannot acquire the cleanup
// write lock. Healthy users must still load; the expired one stays denied.
fresh := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(&iamExpiryLockFailure{ObjectLayer: obj, path: path}, MinIOUsersSysType)}
mustIAM(t, fresh.LoadIAMCache(ctx, true))
if _, ok := fresh.GetUser("healthy-user"); !ok {
t.Fatal("cleanup failure prevented healthy IAM state from loading")
}
if _, ok := fresh.GetUser(c.AccessKey); ok {
t.Fatal("cleanup failure admitted an expired service account")
}
r, err := loadIAMRevision(ctx, fresh, path)
mustIAM(t, err)
if r.Deleted || !r.Credentials.IsExpired() {
t.Fatal("failed cleanup lost the existing expired revision")
}
}
+220
View File
@@ -0,0 +1,220 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"fmt"
"maps"
"strings"
"sync"
"time"
"github.com/minio/minio/internal/auth"
)
// This index is rebuilt by the existing IAM loaders and updated by successful
// storage operations. It avoids a second full IAM walk during every heal pass.
// It is an optimization of the durable records, never a reason to delete them.
// The index contains no secrets or grants.
type iamParentRevision struct {
deleted bool
before time.Time
}
type iamRevisionIndex struct {
mu sync.RWMutex
items map[string]iamRevision
parents map[string]iamParentRevision
floors map[string]time.Time
generation uint64
}
func (idx *iamRevisionIndex) observe(path string, data []byte) {
if !strings.HasPrefix(path, iamConfigPrefix+"/") {
return
}
var r iamRevision
if json.Unmarshal(data, &r) != nil {
return // The caller reports malformed data using its normal decoder.
}
r.Credentials = auth.Credentials{ParentUser: r.Credentials.ParentUser, Expiration: r.Credentials.Expiration}
idx.mu.Lock()
defer idx.mu.Unlock()
if strings.HasPrefix(path, iamConfigUsersPrefix) {
// Keep a compact name-keyed view for the authentication hot path;
// constructing a config path on every S3 request allocates needlessly.
defer func() {
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
if current, ok := idx.items[path]; ok {
if idx.parents == nil {
idx.parents = make(map[string]iamParentRevision)
}
idx.parents[name] = iamParentRevision{deleted: current.Deleted, before: current.RevokedBefore}
} else {
delete(idx.parents, name)
}
}()
}
if floor, ok := idx.floors[path]; ok && r.timestamp().Before(floor) {
return
}
if previous, ok := idx.items[path]; ok {
// A concurrent read that began before a write must not roll it back.
if previous.timestamp().After(r.timestamp()) || (previous.Deleted && !r.Deleted && !r.timestamp().After(previous.timestamp())) {
return
}
if previous.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = previous.RevokedBefore
}
if previous.timestamp().Equal(r.timestamp()) && previous.Deleted == r.Deleted && previous.RevokedBefore.Equal(r.RevokedBefore) {
return
}
}
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
if _, tracked := idx.items[path]; tracked {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
return
}
if !r.Deleted && r.RevokedBefore.IsZero() {
_, tracked := idx.items[path]
_, hasFloor := idx.floors[path]
if tracked || hasFloor {
if idx.floors == nil {
idx.floors = make(map[string]time.Time)
}
idx.floors[path] = r.timestamp()
}
if tracked {
delete(idx.items, path)
idx.generation++
}
return
}
if idx.items == nil {
idx.items = make(map[string]iamRevision)
}
idx.items[path] = r
delete(idx.floors, path)
idx.generation++
}
func (idx *iamRevisionIndex) get(path string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
defer idx.mu.RUnlock()
return idx.items[path]
}
func (idx *iamRevisionIndex) snapshot() map[string]iamRevision {
idx.mu.Lock()
defer idx.mu.Unlock()
for path, r := range idx.items {
if r.Deleted && !r.ExpiresAt.IsZero() && UTCNow().After(r.ExpiresAt) {
delete(idx.items, path)
delete(idx.floors, path)
idx.generation++
}
}
return maps.Clone(idx.items)
}
func (idx *iamRevisionIndex) count() int {
idx.mu.RLock()
defer idx.mu.RUnlock()
return len(idx.items)
}
// A process-local generation plus the protocol's instance ID is sufficient
// for acknowledgements. Avoid hashing the entire index on every IAM write.
func (idx *iamRevisionIndex) digest() string {
idx.mu.RLock()
defer idx.mu.RUnlock()
return fmt.Sprintf("%x:%x", idx.generation, len(idx.items))
}
func (idx *iamRevisionIndex) forget(path string) {
idx.mu.Lock()
if _, ok := idx.items[path]; ok {
delete(idx.items, path)
idx.generation++
}
delete(idx.floors, path)
if strings.HasPrefix(path, iamConfigUsersPrefix) {
delete(idx.parents, strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile))
}
idx.mu.Unlock()
}
func (c *iamCache) userRevocation(user string) iamRevision {
r := c.revisions.parentRevision(user)
if u, ok := c.iamUsersMap[user]; ok && u.RevokedBefore.After(r.RevokedBefore) {
r.RevokedBefore = u.RevokedBefore
}
return r
}
func (c *iamCache) groupMemberAllowed(member string, grantedAt, groupBoundary time.Time) bool {
r := c.userRevocation(member)
return !r.Deleted && (r.RevokedBefore.IsZero() || grantedAt.After(r.RevokedBefore)) && (groupBoundary.IsZero() || grantedAt.After(groupBoundary))
}
func iamMappingParentPath(path string) string {
kind, name, ok := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !ok {
return ""
}
name = strings.TrimSuffix(name, ".json")
switch kind {
case "users", "sts-users":
return getUserIdentityPath(name, regUser)
case "service-accounts":
return getUserIdentityPath(name, svcUser)
case "groups":
return getGroupInfoPath(name)
}
return ""
}
func (idx *iamRevisionIndex) mappingAllowed(path string, mp MappedPolicy) bool {
if mp.Deleted || idx.get(path).Deleted {
return false
}
r := idx.get(iamMappingParentPath(path))
return !r.Deleted && (r.RevokedBefore.IsZero() || mp.UpdatedAt.After(r.RevokedBefore))
}
// Apply the persisted commit boundary even before dependent cache cleanup has
// completed. The map namespace is part of the authorization record's identity.
func (c *iamCache) cachedMappedPolicy(name string, userType IAMUserType, isGroup bool) (MappedPolicy, bool) {
var mp MappedPolicy
var ok bool
switch {
case isGroup:
mp, ok = c.iamGroupPolicyMap.Load(name)
case userType == stsUser:
mp, ok = c.iamSTSPolicyMap.Load(name)
default:
mp, ok = c.iamUserPolicyMap.Load(name)
}
if !ok || !c.revisions.mappingAllowed(getMappedPolicyPath(name, userType, isGroup), mp) {
return MappedPolicy{}, false
}
return mp, true
}
func (idx *iamRevisionIndex) parentRevision(user string) iamRevision {
if idx == nil {
return iamRevision{}
}
idx.mu.RLock()
p := idx.parents[user]
idx.mu.RUnlock()
return iamRevision{Deleted: p.deleted, RevokedBefore: p.before}
}
+305
View File
@@ -0,0 +1,305 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"fmt"
"os"
"strings"
"sync"
"testing"
"time"
"github.com/minio/madmin-go/v3"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
"go.etcd.io/etcd/client/v3/namespace"
)
type iamRevisionLockObserver struct {
ObjectLayer
path string
waiting chan struct{}
once sync.Once
}
func (o *iamRevisionLockObserver) NewNSLock(bucket string, objects ...string) RWLocker {
lock := o.ObjectLayer.NewNSLock(bucket, objects...)
if bucket == minioMetaBucket && len(objects) == 1 && objects[0] == o.path {
return &iamRevisionObservedLock{RWLocker: lock, observe: func() { o.once.Do(func() { close(o.waiting) }) }}
}
return lock
}
type iamRevisionObservedLock struct {
RWLocker
observe func()
}
func (l *iamRevisionObservedLock) GetLock(ctx context.Context, timeout *dynamicTimeout) (LockContext, error) {
l.observe()
return l.RWLocker.GetLock(ctx, timeout)
}
type iamRevisionWatchObserver struct {
etcd.Watcher
waiting chan struct{}
once sync.Once
}
func (w *iamRevisionWatchObserver) Watch(ctx context.Context, key string, opts ...etcd.OpOption) etcd.WatchChan {
w.once.Do(func() { close(w.waiting) })
return w.Watcher.Watch(ctx, key, opts...)
}
// Simulate an unavailable cleanup RPC. Mutex.Lock calls Delete after its wait
// is canceled; that RPC must inherit a deadline too, not Client.Ctx() forever.
type iamRevisionCleanupBlocker struct {
etcd.KV
release chan struct{}
}
func (b *iamRevisionCleanupBlocker) Delete(ctx context.Context, key string, opts ...etcd.OpOption) (*etcd.DeleteResponse, error) {
if strings.Contains(key, "/iam-revision-locks/") {
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-b.release:
}
}
return b.KV.Delete(ctx, key, opts...)
}
type iamRevisionReadBlocker struct {
IAMStorageAPI
path string
after int
waiting chan struct{}
}
func (b *iamRevisionReadBlocker) loadIAMConfig(ctx context.Context, item any, path string) error {
if path == b.path {
b.after--
if b.after == 0 {
close(b.waiting)
<-ctx.Done()
return ctx.Err()
}
}
return b.IAMStorageAPI.loadIAMConfig(ctx, item, path)
}
func TestIAMRevisionReadDoesNotBlockAuthentication(t *testing.T) {
for _, stage := range []struct {
name string
offset time.Duration
}{{"deletion", time.Minute}, {"retained_revocation", -time.Minute}} {
t.Run(stage.name, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
store := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
const user = "read-blocked-parent"
created, err := store.AddUser(ctx, user, madmin.AddOrUpdateUserReq{SecretKey: "original-password", Status: madmin.AccountEnabled})
if err != nil {
t.Fatal(err)
}
blocked := &iamRevisionReadBlocker{IAMStorageAPI: store.IAMStorageAPI, path: getUserIdentityPath(user, regUser), after: 1, waiting: make(chan struct{})}
store.IAMStorageAPI = blocked
done := make(chan error, 1)
go func() {
done <- store.DeleteUser(withIAMReplicationTime(ctx, created.Add(stage.offset)), user, regUser)
}()
defer func() { cancel(); <-done }()
select {
case <-blocked.waiting:
case <-time.After(5 * time.Second):
t.Fatal("revision read was not attempted")
}
read := make(chan bool, 1)
go func() {
u, ok := store.GetUser(user)
read <- ok && u.Credentials.SecretKey == "original-password"
}()
select {
case ok := <-read:
if !ok {
t.Fatal("pending revision read changed the cached identity")
}
case <-time.After(time.Second):
t.Fatal("revision read blocked cached authentication")
}
})
}
}
func TestIAMRevisionLockContention(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if backend == "etcd" && endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
for _, outcome := range []string{"release", "cancel", "default_timeout"} {
t.Run(outcome, func(t *testing.T) {
resetTestGlobals()
t.Cleanup(resetTestGlobals)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
oldTimeout := defaultContextTimeout
defaultContextTimeout = 2 * time.Second
t.Cleanup(func() { defaultContextTimeout = oldTimeout })
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
const user = "contended-user"
path := getUserIdentityPath(user, regUser)
waiting := make(chan struct{})
var store *IAMStoreSys
var hold func() func()
unblockCleanup := func() {}
if backend == "object" {
disks, err := getRandomDisks(1)
must(err)
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
must(err)
t.Cleanup(func() {
obj.Shutdown(context.Background())
os.RemoveAll(disks[0])
})
observed := &iamRevisionLockObserver{ObjectLayer: obj, path: path + ".revision-lock", waiting: waiting}
store = &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, MinIOUsersSysType)}
hold = func() func() {
lock := obj.NewNSLock(minioMetaBucket, observed.path)
lc, err := lock.GetLock(ctx, newDynamicTimeout(time.Second, time.Second))
must(err)
store.IAMStorageAPI.(*IAMObjectStore).objAPI = observed
return func() { lock.Unlock(lc) }
}
} else {
client, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: time.Second})
must(err)
t.Cleanup(func() { client.Close() })
prefix := fmt.Sprintf("/silo-lock-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(client.KV, prefix)
client.Watcher = namespace.NewWatcher(client.Watcher, prefix)
store = &IAMStoreSys{IAMStorageAPI: newIAMEtcdStore(client, MinIOUsersSysType)}
hold = func() func() {
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
must(err)
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
must(lock.Lock(ctx))
client.Watcher = &iamRevisionWatchObserver{Watcher: client.Watcher, waiting: waiting}
blocker := &iamRevisionCleanupBlocker{KV: client.KV, release: make(chan struct{})}
client.KV = blocker
unblockCleanup = sync.OnceFunc(func() { close(blocker.release) })
t.Cleanup(unblockCleanup)
return func() { session.Close() }
}
}
request := func(secret string) madmin.AddOrUpdateUserReq {
return madmin.AddOrUpdateUserReq{SecretKey: secret, Status: madmin.AccountEnabled}
}
_, err := store.AddUser(ctx, user, request("original-password"))
must(err)
release := sync.OnceFunc(hold())
t.Cleanup(release)
writeCtx, cancelWrite := context.WithCancel(ctx)
defer cancelWrite()
first, second := make(chan error, 1), make(chan error, 1)
var writers sync.WaitGroup
t.Cleanup(func() {
cancelWrite()
unblockCleanup()
release()
writers.Wait()
})
writers.Go(func() {
_, err := store.AddUser(writeCtx, user, request("first-password"))
first <- err
})
select {
case <-waiting:
case <-time.After(5 * time.Second):
t.Fatal("writer did not attempt the held revision lock")
}
// A second writer must queue without taking the cache's RWMutex:
// Go's writer preference would otherwise block every new reader.
writers.Go(func() {
_, err := store.AddUser(ctx, user, request("second-password"))
second <- err
})
select {
case err := <-second:
t.Fatalf("second writer bypassed the first: %v", err)
case <-time.After(50 * time.Millisecond):
}
read := make(chan UserIdentity, 1)
go func() {
u, _ := store.GetUser(user)
read <- u
}()
select {
case u := <-read:
if u.Credentials.SecretKey != "original-password" {
t.Fatal("pending write changed the cached credential")
}
case <-time.After(time.Second):
t.Fatal("distributed lock contention blocked cached authentication")
}
switch outcome {
case "release":
release()
case "cancel":
cancelWrite()
}
select {
case err := <-first:
if outcome == "release" {
must(err)
} else if err == nil {
t.Fatal("canceled or timed-out write succeeded")
}
case <-time.After(5 * time.Second):
t.Fatal("lock wait or cancellation cleanup exceeded its deadline")
}
release()
select {
case err := <-second:
must(err)
case <-time.After(5 * time.Second):
t.Fatal("queued writer did not recover after the first completed")
}
cached, ok := store.GetUser(user)
if !ok || cached.Credentials.SecretKey != "second-password" {
t.Fatal("cached write order was lost")
}
var persisted UserIdentity
must(store.loadIAMConfig(ctx, &persisted, path))
if persisted.Credentials.SecretKey != cached.Credentials.SecretKey || !persisted.UpdatedAt.Equal(cached.UpdatedAt) {
t.Fatal("persistent and cached revisions differ")
}
})
}
})
}
}
+749
View File
@@ -0,0 +1,749 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"crypto/sha256"
"errors"
"fmt"
"net/http"
"sort"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/concurrency"
)
var errIAMStaleUpdate = errors.New("IAM update predates a stored revision or revocation")
// The parent is still live; callers must not broadcast a user deletion when
// only its revocation boundary was retained.
var errIAMRevocationRetained = errors.New("IAM revocation recorded without deleting the record")
// A revocation advances the boundary even when a newer identity already
// exists. Keep this operation distinct from replacing/deleting that identity.
type iamUserRevocation struct {
UserIdentity
retained bool
}
type iamGroupRevocation struct {
GroupInfo
retained bool
requireEmpty bool
}
// Natural expiration is distinct from revoking a live credential. An expired
// immutable STS token can be removed; a reusable service-account key retains
// its revision so an older non-expiring credential cannot return.
type iamExpireIdentity struct{}
// The authoritative revocation is durable even if dependent cleanup fails.
// Callers must publish it to sibling caches before returning the error.
type iamCommittedCleanupError struct {
err error
retained bool
}
func (e *iamCommittedCleanupError) Error() string {
return "IAM revocation committed; cleanup failed: " + e.err.Error()
}
func (e *iamCommittedCleanupError) Unwrap() error { return e.err }
type iamReplicationTimeKey struct{}
func withIAMReplicationTime(ctx context.Context, at time.Time) context.Context {
return context.WithValue(ctx, iamReplicationTimeKey{}, at)
}
func iamReplicationTime(ctx context.Context) (time.Time, bool) {
at, ok := ctx.Value(iamReplicationTimeKey{}).(time.Time)
return at, ok
}
func iamReplicationError(err error) error {
if errors.Is(err, errIAMStaleUpdate) {
// Retrying an obsolete event cannot change the result.
return nil
}
return wrapSRErr(err)
}
// Deletions occupy the original IAM config path. They contain no secret or
// grant and are hidden by the normal loaders, but remain available to heal
// and to timestamp comparisons after a restart. Do not age them out: a peer
// can be offline indefinitely.
type iamRevision struct {
UpdatedAt time.Time `json:"updatedAt"`
UpdateDate time.Time `json:"UpdateDate"`
Deleted bool `json:"deleted"`
RevokedBefore time.Time `json:"revokedBefore"`
ExpiresAt time.Time `json:"expiresAt,omitempty"`
Credentials auth.Credentials `json:"credentials"`
}
func (r iamRevision) timestamp() time.Time {
if r.UpdateDate.After(r.UpdatedAt) {
return r.UpdateDate
}
return r.UpdatedAt
}
func loadIAMRevision(ctx context.Context, store IAMStorageAPI, path string) (iamRevision, error) {
var r iamRevision
err := store.loadIAMConfig(ctx, &r, path)
if errors.Is(err, errConfigNotFound) {
err = nil
}
return r, err
}
func (store *IAMStoreSys) checkIAMRevision(ctx context.Context, path string, deleting bool) error {
at, replicated := iamReplicationTime(ctx)
if !replicated {
return nil
}
return store.withIAMStorage(ctx, func(ctx context.Context) error {
r, err := loadIAMRevision(ctx, store.IAMStorageAPI, path)
if err != nil {
return err
}
if r.timestamp().After(at) || (r.Deleted && !deleting && !at.After(r.timestamp())) {
return errIAMStaleUpdate
}
return nil
})
}
// This signed claim records the parent's revocation boundary at issuance.
// Unlike UpdatedAt, it cannot advance when an offline site edits an old child.
// It travels in the existing service-account Claims and STS SessionToken fields.
const iamParentRevocationClaim = "siloParentRevocation"
func setIAMParentRevocationClaim(ctx context.Context, store IAMStorageAPI, parent string, claims map[string]any) error {
delete(claims, iamParentRevocationClaim)
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted {
return errIAMStaleUpdate
}
if !r.RevokedBefore.IsZero() {
claims[iamParentRevocationClaim] = r.RevokedBefore.Format(time.RFC3339Nano)
}
return nil
}
func iamCredentialSurvivesRevocation(cred auth.Credentials, at time.Time) bool {
if at.IsZero() {
return true
}
s, _ := cred.Claims[iamParentRevocationClaim].(string)
issuedAfter, err := time.Parse(time.RFC3339Nano, s)
return err == nil && !issuedAfter.Before(at)
}
// Parent revocations delete old children even if an offline peer has edited
// them later. Preserve children that prove issuance after this revocation.
func iamChildDeletionContext(ctx context.Context, child UserIdentity) (context.Context, bool) {
if at, replicated := iamReplicationTime(ctx); replicated {
if !at.IsZero() && iamCredentialSurvivesRevocation(child.Credentials, at) {
return ctx, false
}
if child.UpdatedAt.After(at) {
ctx = withIAMReplicationTime(ctx, child.UpdatedAt)
}
}
return ctx, true
}
// A delayed service account or STS event must not outlive deletion of its
// built-in parent. The caller must populate Claims from the verified token.
func checkIAMParentRevision(ctx context.Context, store IAMStorageAPI, cred auth.Credentials) error {
parent := cred.ParentUser
if parent == "" || parent == globalActiveCred.AccessKey {
return nil
}
r, err := loadIAMRevision(ctx, store, getUserIdentityPath(parent, regUser))
if err != nil {
return err
}
if r.Deleted || !iamCredentialSurvivesRevocation(cred, r.RevokedBefore) {
return errIAMStaleUpdate
}
return nil
}
// Called with the IAM writer mutex and cache lock held. Persistence only
// touches the caller's record, not the cache. Keep writers serialized while
// allowing cached authentication reads throughout storage and lock waits.
func (store *IAMStoreSys) withIAMStorage(ctx context.Context, fn func(context.Context) error) error {
store.IAMStorageAPI.unlock()
defer store.IAMStorageAPI.lock()
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
return fn(ctx)
}
func (store *IAMStoreSys) saveIAMRevision(ctx context.Context, path string, item any, opts ...options) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return saveIAMRevision(ctx, store.IAMStorageAPI, path, item, opts...)
})
}
func (store *IAMStoreSys) checkIAMParentRevision(ctx context.Context, cred auth.Credentials) error {
return store.withIAMStorage(ctx, func(ctx context.Context) error {
return checkIAMParentRevision(ctx, store.IAMStorageAPI, cred)
})
}
// Update the caller's record with the persisted revision before it is cached.
func saveIAMRevision(ctx context.Context, store IAMStorageAPI, path string, item any, opts ...options) error {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
// Serialize compare-and-write across nodes, as well as goroutines. Use a
// separate lock name so saving the config does not reacquire this lock.
switch s := store.(type) {
case *IAMObjectStore:
lock := s.objAPI.NewNSLock(minioMetaBucket, path+".revision-lock")
lc, err := lock.GetLock(ctx, globalOperationTimeout)
if err != nil {
return err
}
defer lock.Unlock(lc)
ctx = lc.Context()
case *IAMEtcdStore:
// Mutex.Lock also uses Client.Ctx() for cleanup after cancellation.
// Borrow the existing services with the operation's bounded context;
// never close this facade, which does not own those services.
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(s.client.GetLogger()))
client.KV, client.Lease, client.Watcher = s.client.KV, s.client.Lease, s.client.Watcher
session, err := concurrency.NewSession(client, concurrency.WithContext(ctx))
if err != nil {
return err
}
defer func() {
session.Orphan()
// A canceled operation must still release its lease when etcd is
// reachable. If it is unavailable, stop waiting and let it expire.
cleanupCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), defaultContextTimeout)
defer cancel()
_, _ = s.client.Revoke(cleanupCtx, session.Lease())
}()
lock := concurrency.NewMutex(session, fmt.Sprintf("%s/iam-revision-locks/%x", minioConfigPrefix, sha256.Sum256([]byte(path))))
if err = lock.Lock(ctx); err != nil {
return err
}
// Revoking the session lease releases the lock, including on cancellation.
}
previous, err := loadIAMRevision(ctx, store, path)
if err != nil {
return err
}
if _, expiring := item.(*iamExpireIdentity); expiring {
sts := strings.HasPrefix(path, iamConfigSTSPrefix)
if previous.Deleted {
if sts && !previous.ExpiresAt.IsZero() && UTCNow().After(previous.ExpiresAt) {
return expireIAMSTSConfig(ctx, store, path)
}
return nil
}
if previous.timestamp().IsZero() || !previous.Credentials.IsExpired() {
return nil
}
if sts {
return expireIAMSTSConfig(ctx, store, path)
}
item = &UserIdentity{Version: 1, Deleted: true}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
var revocation *iamUserRevocation
if op, ok := item.(*iamUserRevocation); ok {
revocation = op
op.UserIdentity = UserIdentity{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.UserIdentity = UserIdentity{Version: 1, Credentials: previous.Credentials, UpdatedAt: previous.timestamp(), RevokedBefore: origin}
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.UserIdentity
}
var groupRevocation *iamGroupRevocation
if op, ok := item.(*iamGroupRevocation); ok {
groupRevocation = op
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
if op.requireEmpty && !group.Deleted {
for _, member := range group.Members {
r := store.revisionIndex().get(getUserIdentityPath(member, regUser))
at := group.MemberGrants[member]
if !r.Deleted && (r.RevokedBefore.IsZero() || at.After(r.RevokedBefore)) && (group.RevokedBefore.IsZero() || at.After(group.RevokedBefore)) {
return errGroupNotEmpty
}
}
}
op.GroupInfo = GroupInfo{Version: 1, Deleted: true}
if origin, replicated := iamReplicationTime(ctx); replicated && previous.timestamp().After(origin) {
if previous.Deleted || !origin.After(previous.RevokedBefore) {
return errIAMStaleUpdate
}
op.retained = true
op.GroupInfo = group
op.RevokedBefore = origin
ctx = withIAMReplicationTime(ctx, previous.timestamp())
}
item = &op.GroupInfo
}
var at *time.Time
var deleted bool
switch v := item.(type) {
case *UserIdentity:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *GroupInfo:
at, deleted = &v.UpdatedAt, v.Deleted
if boundary, ok := ctx.Value(iamRecordBoundaryKey{}).(time.Time); ok && boundary.After(v.RevokedBefore) {
v.RevokedBefore = boundary
}
if !deleted {
var group GroupInfo
if err := store.loadIAMConfig(ctx, &group, path); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
mergeIAMGroupMutation(ctx, group, v)
}
if previous.RevokedBefore.After(v.RevokedBefore) {
v.RevokedBefore = previous.RevokedBefore
}
case *MappedPolicy:
at, deleted = &v.UpdatedAt, v.Deleted
case *PolicyDoc:
at, deleted = &v.UpdateDate, v.Deleted
default:
return errInvalidArgument
}
if strings.HasPrefix(path, iamConfigSTSPrefix) && previous.Deleted && !deleted {
// STS access keys identify immutable tokens, not reusable user names.
return errIAMStaleUpdate
}
if origin, replicated := iamReplicationTime(ctx); replicated {
*at = origin
if previous.timestamp().After(origin) || (previous.Deleted && !deleted && !origin.After(previous.timestamp())) {
return errIAMStaleUpdate
}
if strings.HasPrefix(path, iamConfigServiceAccountsPrefix) && !deleted && previous.Credentials.AccessKey != "" && previous.timestamp().Equal(origin) {
// Duplicate service snapshots are acknowledgements, not new creates
// or edits. Reload the winner without writing, so even a stale
// sibling cache is refreshed by the retry before acknowledging it.
return store.loadIAMConfig(ctx, item, path)
}
if previous.Deleted && deleted && !origin.After(previous.timestamp()) {
// An already-applied tombstone needs no further persistent write.
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
} else {
if previous.Deleted && deleted {
// A peer notification without an originating revision must not
// advance a tombstone past a subsequent deliberate recreation.
*at = previous.timestamp()
if v, ok := item.(*UserIdentity); ok {
v.RevokedBefore = previous.RevokedBefore
}
if v, ok := item.(*GroupInfo); ok {
v.RevokedBefore = previous.RevokedBefore
}
return nil
}
if at.IsZero() {
*at = UTCNow()
}
if !at.After(previous.timestamp()) {
*at = previous.timestamp().Add(time.Nanosecond)
}
}
if v, ok := item.(*UserIdentity); ok {
if deleted {
// Retain only the parent name for root-account exclusion during heal.
v.Credentials = auth.Credentials{ParentUser: previous.Credentials.ParentUser}
v.RevokedBefore = *at
if strings.HasPrefix(path, iamConfigSTSPrefix) && !previous.Credentials.Expiration.IsZero() && !previous.Credentials.Expiration.Equal(timeSentinel) {
// The signed STS token cannot authorize beyond this time, even
// if an offline site replays it with a newer event timestamp.
v.ExpiresAt = previous.Credentials.Expiration.Add(globalMaxSkewTime)
opts = []options{{ttl: max(1, int64(time.Until(v.ExpiresAt).Seconds())+1)}}
}
} else {
if v.Credentials.SessionToken != "" && v.Credentials.Claims == nil {
claims, err := extractJWTClaims(*v)
if err != nil {
return err
}
v.Credentials.Claims = claims.Map()
}
if err = checkIAMParentRevision(ctx, store, v.Credentials); err != nil {
return err
}
}
}
if v, ok := item.(*GroupInfo); ok && deleted {
v.RevokedBefore = *at
v.Members, v.MemberGrants = nil, nil
}
if _, ok := item.(*MappedPolicy); ok && !deleted {
if parentPath := iamMappingParentPath(path); parentPath != "" {
parent, err := loadIAMRevision(ctx, store, parentPath)
if err != nil {
return err
}
if parent.Deleted {
return errIAMStaleUpdate
}
if !parent.RevokedBefore.IsZero() && !at.After(parent.RevokedBefore) {
if _, replicated := iamReplicationTime(ctx); replicated {
return errIAMStaleUpdate
}
*at = parent.RevokedBefore.Add(time.Nanosecond)
}
}
}
if err := store.saveIAMConfig(ctx, item, path, opts...); err != nil {
return err
}
if revocation != nil && revocation.retained {
return errIAMRevocationRetained
}
if groupRevocation != nil && groupRevocation.retained {
return errIAMRevocationRetained
}
return nil
}
func (iamOS *IAMObjectStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithCancel(ctx)
defer cancel()
var paths []string
for item := range listIAMConfigItems(ctx, iamOS.objAPI, iamConfigPrefix+"/") {
if item.Err != nil {
return nil, item.Err
}
paths = append(paths, iamConfigPrefix+"/"+item.Item)
}
return paths, nil
}
func (ies *IAMEtcdStore) listIAMConfigPaths(ctx context.Context) ([]string, error) {
ctx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
r, err := ies.client.Get(ctx, iamConfigPrefix+"/", etcd.WithPrefix(), etcd.WithKeysOnly())
if err != nil {
return nil, err
}
paths := make([]string, 0, len(r.Kvs))
for _, kv := range r.Kvs {
paths = append(paths, string(kv.Key))
}
return paths, nil
}
func iamDeletionItem(path string, r iamRevision) (item madmin.SRIAMItem, ok bool) {
if (strings.HasPrefix(path, iamConfigUsersPrefix) || strings.HasPrefix(path, iamConfigGroupsPrefix)) && !r.RevokedBefore.IsZero() {
// Recreating a parent does not cancel its older revocation of derived
// credentials. Replay this boundary even after the parent is live again.
r.Deleted = true
r.UpdatedAt, r.UpdateDate = r.RevokedBefore, time.Time{}
}
if !r.Deleted {
return item, false
}
item.UpdatedAt = r.timestamp()
switch {
case strings.HasPrefix(path, iamConfigUsersPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigUsersPrefix), "/"+iamIdentityFile)
item.Type = madmin.SRIAMItemIAMUser
item.IAMUser = &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}
case strings.HasPrefix(path, iamConfigServiceAccountsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigServiceAccountsPrefix), "/"+iamIdentityFile)
if name == siteReplicatorSvcAcc || r.Credentials.ParentUser == globalActiveCred.AccessKey {
return item, false
}
item.Type = madmin.SRIAMItemSvcAcc
item.SvcAccChange = &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: name}}
case strings.HasPrefix(path, iamConfigGroupsPrefix):
name := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigGroupsPrefix), "/"+iamGroupMembersFile)
item.Type = madmin.SRIAMItemGroupInfo
item.GroupInfo = &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: name, IsRemove: true}}
case strings.HasPrefix(path, iamConfigPoliciesPrefix):
item.Type = madmin.SRIAMItemPolicy
item.Name = strings.TrimSuffix(strings.TrimPrefix(path, iamConfigPoliciesPrefix), "/"+iamPolicyFile)
case strings.HasPrefix(path, iamConfigPolicyDBPrefix):
prefix, name, found := strings.Cut(strings.TrimPrefix(path, iamConfigPolicyDBPrefix), "/")
if !found {
return item, false
}
typ := regUser
switch prefix {
case "sts-users":
typ = stsUser
case "service-accounts":
typ = svcUser
}
item.Type = madmin.SRIAMItemPolicyMapping
item.PolicyMapping = &madmin.SRPolicyMapping{UserOrGroup: strings.TrimSuffix(name, ".json"), UserType: int(typ), IsGroup: prefix == "groups"}
default:
// Expired STS credentials are not replayed. Parent revocations and
// their retained timestamp reject delayed copies of derived tokens.
return item, false
}
return item, true
}
func iamDeletionPath(item madmin.SRIAMItem) string {
switch item.Type {
case madmin.SRIAMItemIAMUser:
if item.IAMUser != nil && item.IAMUser.IsDeleteReq {
return getUserIdentityPath(item.IAMUser.AccessKey, regUser)
}
case madmin.SRIAMItemSvcAcc:
if item.SvcAccChange != nil && item.SvcAccChange.Delete != nil {
return getUserIdentityPath(item.SvcAccChange.Delete.AccessKey, svcUser)
}
case madmin.SRIAMItemGroupInfo:
if item.GroupInfo != nil && item.GroupInfo.UpdateReq.IsRemove && len(item.GroupInfo.UpdateReq.Members) == 0 {
return getGroupInfoPath(item.GroupInfo.UpdateReq.Group)
}
case madmin.SRIAMItemPolicy:
if len(item.Policy) == 0 {
return getPolicyDocPath(item.Name)
}
case madmin.SRIAMItemPolicyMapping:
if p := item.PolicyMapping; p != nil && p.Policy == "" {
return getMappedPolicyPath(p.UserOrGroup, IAMUserType(p.UserType), p.IsGroup)
}
}
return ""
}
func (c *SiteReplicationSys) healIAMDeletions(ctx context.Context) (err error) {
started := time.Now()
defer func() {
c.iamRevisionMetrics.healDurationMillis.Store(time.Since(started).Milliseconds())
if err != nil {
c.iamRevisionMetrics.healFailures.Add(1)
} else {
c.iamRevisionMetrics.healLastSuccess.Store(time.Now().Unix())
}
}()
c.iamHealMu.Lock()
defer c.iamHealMu.Unlock()
c.RLock()
defer c.RUnlock()
if !c.enabled {
return nil
}
snapshot := globalIAMSys.store.revisionIndex().snapshot()
paths := make([]string, 0, len(snapshot))
for path := range snapshot {
paths = append(paths, path)
}
sort.Strings(paths)
byType := make(map[string][]iamReplicationItem)
for _, path := range paths {
r := snapshot[path]
item, ok := iamDeletionItem(path, r)
if !ok {
continue
}
out := iamReplicationItem{SRIAMItem: item}
if item.Type == madmin.SRIAMItemIAMUser && !r.Deleted {
out.Type, out.IAMUser = iamUserBoundaryType, nil
out.UserRevocation = &iamUserBoundary{User: item.IAMUser.AccessKey, Before: r.RevokedBefore}
}
if item.Type == madmin.SRIAMItemGroupInfo && !r.Deleted {
out.Type, out.GroupInfo = iamGroupBoundaryType, nil
out.GroupRevocation = &iamGroupBoundary{Group: item.GroupInfo.UpdateReq.Group, Before: r.RevokedBefore}
}
byType[item.Type] = append(byType[item.Type], out)
}
var items []iamReplicationItem
for _, typ := range []string{madmin.SRIAMItemPolicyMapping, madmin.SRIAMItemIAMUser, madmin.SRIAMItemSvcAcc, madmin.SRIAMItemGroupInfo, madmin.SRIAMItemPolicy} {
items = append(items, byType[typ]...)
}
if len(items) == 0 {
return nil
}
if c.iamRevisionProgress == nil {
c.iamRevisionProgress = make(map[string]iamRevisionProgress)
}
for id := range c.iamRevisionProgress {
if _, present := c.state.Peers[id]; !present {
delete(c.iamRevisionProgress, id)
}
}
var progressMu sync.Mutex
cerr := c.concDo(nil, func(id string, p madmin.PeerInfo) error {
// Bound each pass, but retain acknowledgements independently of the
// pass deadline or unrelated changes at either site.
peerCtx, cancel := context.WithTimeout(ctx, defaultContextTimeout)
defer cancel()
client, err := c.getAdminClient(peerCtx, id)
if err != nil {
return err
}
remote, err := executeIAMRevisionRequest(peerCtx, client, http.MethodGet, nil)
if err != nil {
return err
}
progressMu.Lock()
progress := c.iamRevisionProgress[id]
progressMu.Unlock()
progress.observePeer(remote)
defer func() {
progressMu.Lock()
c.iamRevisionProgress[id] = progress
progressMu.Unlock()
}()
var pending []iamReplicationItem
for _, item := range items {
path, version := iamReplicationMarker(item)
if progress.Acknowledged[path] != version {
pending = append(pending, item)
}
}
// Acknowledgements are only a replay optimization, never GC proof.
for path := range progress.Acknowledged {
if _, retained := snapshot[path]; !retained {
delete(progress.Acknowledged, path)
}
}
var failures []error
for next := 0; next < len(pending); {
end := min(next+maxIAMRevisionBatch, len(pending))
batch := pending[next:end]
remote, err = executeIAMRevisionRequest(peerCtx, client, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: batch})
if err != nil {
var batchErr *iamRevisionBatchError
if !errors.As(err, &batchErr) {
return errors.Join(append(failures, err)...)
}
failures = append(failures, err)
} else {
progress.observePeer(remote)
for _, item := range batch {
path, version := iamReplicationMarker(item)
progress.Acknowledged[path] = version
}
}
next = end
}
return errors.Join(failures...)
}, "IAM revision convergence")
return errors.Unwrap(cerr)
}
func iamReplicationMarker(item iamReplicationItem) (path, version string) {
path = iamDeletionPath(item.SRIAMItem)
if item.UserRevocation != nil {
path = getUserIdentityPath(item.UserRevocation.User, regUser)
}
if item.GroupRevocation != nil {
path = getGroupInfoPath(item.GroupRevocation.Group)
}
return path, item.Type + ":" + item.UpdatedAt.UTC().Format(time.RFC3339Nano)
}
func (store *IAMStoreSys) savePolicyDoc(ctx context.Context, policyName string, p *PolicyDoc) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(policyName), p)
}
func (store *IAMStoreSys) saveMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool, mp *MappedPolicy, opts ...options) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), mp, opts...)
}
func (store *IAMStoreSys) saveUserIdentity(ctx context.Context, name string, userType IAMUserType, u *UserIdentity, opts ...options) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), u, opts...)
}
func (store *IAMStoreSys) saveGroupInfo(ctx context.Context, name string, gi *GroupInfo) error {
return store.saveIAMRevision(ctx, getGroupInfoPath(name), gi)
}
func (store *IAMStoreSys) deletePolicyDoc(ctx context.Context, name string) error {
return store.saveIAMRevision(ctx, getPolicyDocPath(name), &PolicyDoc{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteMappedPolicy(ctx context.Context, name string, userType IAMUserType, isGroup bool) error {
return store.saveIAMRevision(ctx, getMappedPolicyPath(name, userType, isGroup), &MappedPolicy{Version: 1, Deleted: true})
}
func (store *IAMStoreSys) deleteUserIdentity(ctx context.Context, name string, userType IAMUserType) error {
return store.saveIAMRevision(ctx, getUserIdentityPath(name, userType), &UserIdentity{Version: 1, Deleted: true})
}
// Called under the identity's distributed revision lock, after verifying that
// its immutable STS token (or early-revocation retention) has expired. Only the
// old token-key mapping is removed; the reusable parent mapping is unaffected.
func expireIAMSTSConfig(ctx context.Context, store IAMStorageAPI, path string) error {
key := strings.TrimSuffix(strings.TrimPrefix(path, iamConfigSTSPrefix), "/"+iamIdentityFile)
if err := store.deleteIAMConfig(ctx, getMappedPolicyPath(key, stsUser, false)); err != nil && !errors.Is(err, errConfigNotFound) {
return err
}
return store.deleteIAMConfig(ctx, path)
}
type (
iamExpirationCleanupKey struct{}
iamExpirationCleanupState struct{ failed atomic.Bool }
)
func withIAMExpirationCleanup(ctx context.Context) context.Context {
if _, ok := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState); ok {
return ctx
}
return context.WithValue(ctx, iamExpirationCleanupKey{}, &iamExpirationCleanupState{})
}
func bestEffortIAMExpiration(ctx context.Context, store IAMStorageAPI, path string) {
state, _ := ctx.Value(iamExpirationCleanupKey{}).(*iamExpirationCleanupState)
if ctx.Err() != nil || (state != nil && state.failed.Load()) {
return
}
ctx, cancel := context.WithTimeout(ctx, time.Second)
defer cancel()
// Failure leaves the expired record and its existing version intact.
// Stop optional reclamation for this load, while still loading healthy
// users. Healthy cleanup has no per-scan quota that could build a backlog.
if err := saveIAMRevision(ctx, store, path, &iamExpireIdentity{}); err != nil {
if state != nil {
state.failed.Store(true)
}
iamLogIf(ctx, err)
}
}
+330
View File
@@ -0,0 +1,330 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/grid"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
// Count physical saves: comparing timestamps alone would miss identical
// tombstones being rewritten on every heal pass.
type iamRevisionWriteCounter struct {
IAMStorageAPI
data []byte
writes int
}
func (s *iamRevisionWriteCounter) loadIAMConfig(_ context.Context, item any, _ string) error {
return json.Unmarshal(s.data, item)
}
func (s *iamRevisionWriteCounter) saveIAMConfig(_ context.Context, item any, _ string, _ ...options) error {
data, err := json.Marshal(item)
if err == nil {
s.data = data
s.writes++
}
return err
}
func TestIAMRevocationTombstoneReplayIsIdempotent(t *testing.T) {
at := time.Date(2026, 9, 14, 12, 0, 0, 0, time.UTC)
for _, record := range []struct {
name string
new func(bool) any
}{
{"user", func(deleted bool) any { return &UserIdentity{Version: 1, Deleted: deleted} }},
{"group", func(deleted bool) any { return &GroupInfo{Version: 1, Deleted: deleted} }},
{"policy", func(deleted bool) any { return &PolicyDoc{Version: 1, Deleted: deleted} }},
{"mapping", func(deleted bool) any { return &MappedPolicy{Version: 1, Deleted: deleted} }},
} {
t.Run(record.name, func(t *testing.T) {
data, err := json.Marshal(iamRevision{Deleted: true, UpdatedAt: at, RevokedBefore: at})
if err != nil {
t.Fatal(err)
}
store := &iamRevisionWriteCounter{data: data}
ctx := context.Background()
for range 3 {
// Site heal carries the original timestamp. Sibling notifications
// have no timestamp; both must leave an applied deletion untouched.
for _, replay := range []context.Context{withIAMReplicationTime(ctx, at), ctx} {
if err := saveIAMRevision(replay, store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
}
}
if store.writes != 0 || string(store.data) != string(data) {
t.Fatalf("replayed tombstone changed storage: writes=%d, record=%s", store.writes, store.data)
}
for _, deleted := range []bool{false, true} {
err := saveIAMRevision(withIAMReplicationTime(ctx, at.Add(-time.Second)), store, record.name, record.new(deleted))
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("older event accepted, deleted=%t: %v", deleted, err)
}
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, at), store, record.name, record.new(false)); !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("equal-time recreation accepted: %v", err)
}
newer := at.Add(time.Minute)
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer), store, record.name, record.new(true)); err != nil {
t.Fatal(err)
}
r, err := loadIAMRevision(ctx, store, record.name)
if err != nil || store.writes != 1 || !r.timestamp().Equal(newer) || !r.Deleted {
t.Fatalf("newer deletion did not advance storage: writes=%d, revision=%+v, error=%v", store.writes, r, err)
}
if err := saveIAMRevision(withIAMReplicationTime(ctx, newer.Add(time.Minute)), store, record.name, record.new(false)); err != nil {
t.Fatalf("newer recreation rejected: %v", err)
}
})
}
}
// Run with both object storage and etcd through TestIAMRevocation*Lifecycle.
// The object-store case uses the real peer RPC and deletion handler, so a
// spurious notification actually destroys the parent instead of only counting it.
func testIAMRevocationReplayAfterRecreation(ctx context.Context, t *testing.T, sys *IAMSys) {
t.Helper()
peer := &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "heal-recreated-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-recreated-child", secretKey: "valid-service-password",
})
must(err)
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(req *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(req)
}))
// Future user updates still use the normal peer reload notification.
must(loadUserRPC.Register(tg.Managers[1], server.LoadUserHandler))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
previousNotifications := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection {
return tg.Managers[0].Connection(tg.Hosts[1])
},
}}}
t.Cleanup(func() { globalNotificationSys = previousNotifications })
assertLive := func(key string) {
t.Helper()
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatalf("live credential %s lost during deletion replay", key)
}
}
assertNoDelete := func() {
t.Helper()
if n := deletes.Load(); n != 0 {
t.Fatalf("retained revocation sent %d destructive sibling notifications", n)
}
}
for range 3 {
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
item, ok := iamDeletionItem(getUserIdentityPath(user, regUser), r)
if !ok || item.IAMUser == nil || !item.UpdatedAt.Equal(deleted) {
t.Fatal("recreated user lost its durable revocation replay")
}
must(peer.PeerIAMUserChangeHandler(ctx, item.IAMUser, item.UpdatedAt))
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
assertLive(child.AccessKey)
assertNoDelete()
}
// A divergent site sends a previously unseen revocation between our old
// boundary and recreation. Retain it and revoke old children, but never
// turn it into an unversioned delete of the recreated parent.
delayed := deleted.Add(time.Minute)
revoke(delayed)
assertNoDelete()
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
if _, ok := sys.GetUser(ctx, child.AccessKey); ok {
t.Fatal("child from before the delayed revocation remains usable")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if r.Deleted || !r.RevokedBefore.Equal(delayed) || !r.timestamp().Equal(recreated) {
t.Fatalf("retained revocation damaged the recreated identity: %+v", r)
}
// A genuinely newer deletion must still reach siblings and remove the
// parent plus credentials issued under its latest revocation boundary.
fresh, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "heal-fresh-child", secretKey: "valid-service-password",
})
must(err)
latest := recreated.Add(time.Minute)
revoke(latest)
wantDeletes := int32(1)
if sys.HasWatcher() {
wantDeletes = 0
}
if n := deletes.Load(); n != wantDeletes {
t.Fatalf("new deletion notifications=%d, want %d", n, wantDeletes)
}
for _, key := range []string{user, fresh.AccessKey} {
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatalf("newer deletion left credential %s usable", key)
}
}
// Exercise the actual sibling handler again against the already persisted
// tombstone. Its context has no revision; it must not re-stamp the record.
_, remoteErr := server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
r, err = loadIAMRevision(ctx, sys.store, getUserIdentityPath(user, regUser))
must(err)
if !r.Deleted || !r.timestamp().Equal(latest) {
t.Fatalf("sibling re-stamped the tombstone: got %s, want %s", r.timestamp(), latest)
}
create(latest.Add(time.Minute))
_, remoteErr = server.DeleteUserHandler(grid.NewMSSWith(map[string]string{peerRESTUser: user}))
if remoteErr != nil {
t.Fatal(remoteErr)
}
assertLive(user)
revoke(latest)
must(sys.store.LoadIAMCache(ctx, false))
assertLive(user)
}
// Counts what a retained revocation actually sends to sibling nodes.
func TestIAMRevocationRetainedReloadsSibling(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err != nil {
t.Fatal(err)
}
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, nil, 2*time.Second)
defer os.RemoveAll(disks[0])
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
user := "retained-parent"
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
deleted, recreated := origin.Add(time.Minute), origin.Add(3*time.Minute)
create := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, at))
}
revoke := func(at time.Time) {
must(peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, at))
}
create(origin)
revoke(deleted)
create(recreated)
child, _, err := sys.NewServiceAccount(ctx, user, nil, newServiceAccountOpts{
accessKey: "retained-child", secretKey: "valid-service-password",
})
must(err)
// A sibling shares persistent state but has an independent IAM cache.
sibling := &IAMStoreSys{IAMStorageAPI: newIAMObjectStore(obj, sys.usersSysType)}
must(sibling.LoadIAMCache(ctx, false))
if _, ok := sibling.GetUser(child.AccessKey); !ok {
t.Fatal("sibling fixture did not load child")
}
tg, err := grid.SetupTestGrid(2)
must(err)
t.Cleanup(tg.Cleanup)
var deletes, loads atomic.Int32
server := &peerRESTServer{}
must(deleteUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
deletes.Add(1)
return server.DeleteUserHandler(r)
}))
must(loadUserRPC.Register(tg.Managers[1], func(r *grid.MSS) (grid.NoPayload, *grid.RemoteErr) {
loads.Add(1)
// LoadUserHandler delegates to this same cache reload method.
if err := sibling.UserNotificationHandler(ctx, r.Get(peerRESTUser), regUser); err != nil {
return grid.NoPayload{}, grid.NewRemoteErr(err)
}
return grid.NoPayload{}, nil
}))
host, err := xnet.ParseHost(strings.TrimPrefix(tg.Hosts[1], "http://"))
must(err)
prev := globalNotificationSys
globalNotificationSys = &NotificationSys{peerClients: []*peerRESTClient{{
host: host,
gridConn: func() *grid.Connection { return tg.Managers[0].Connection(tg.Hosts[1]) },
}}}
t.Cleanup(func() { globalNotificationSys = prev })
delayed := deleted.Add(time.Minute)
revoke(delayed)
t.Logf("sibling notifications after a retained revocation: destructive=%d reload=%d", deletes.Load(), loads.Load())
if deletes.Load() != 0 {
t.Errorf("destructive sibling delete sent: %d", deletes.Load())
}
if loads.Load() == 0 {
t.Errorf("retained revocation did not notify the sibling")
}
if _, ok := sibling.GetUser(child.AccessKey); ok {
t.Error("sibling still resolves revoked child")
}
if _, ok := sibling.GetUser(user); !ok {
t.Error("sibling lost live parent")
}
if _, ok := sys.store.GetUser(user); !ok {
t.Error("live parent lost")
}
if _, ok := sys.store.GetUser(child.AccessKey); ok {
t.Error("revoked child still resolves on the receiving node")
}
}
+535
View File
@@ -0,0 +1,535 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"net/http/httptest"
"os"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
etcd "go.etcd.io/etcd/client/v3"
"go.etcd.io/etcd/client/v3/namespace"
)
// Exercise the persisted IAM store and the same peer handler used by site heal.
// A delete must survive a cache reload and an older create arriving afterwards.
func TestIAMRevocationRejectsOfflineUser(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
user := "offline-revoked-user"
req := madmin.AddOrUpdateUserReq{SecretKey: "test-password-valid", Status: madmin.AccountEnabled}
created, err := globalIAMSys.CreateUser(ctx, user, req)
if err != nil {
t.Fatal(err)
}
if err = globalIAMSys.DeleteUser(ctx, user, false); err != nil {
t.Fatal(err)
}
if err = globalIAMSys.store.LoadIAMCache(ctx, false); err != nil {
t.Fatal(err)
}
if err = globalSiteReplicationSys.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, UserReq: &req}, created); err != nil {
t.Fatal(err)
}
if _, err = globalIAMSys.GetUserInfo(ctx, user); !errors.Is(err, errNoSuchUser) {
t.Fatalf("revoked user restored by old peer event: %v", err)
}
}
func TestIAMRevocationHealingContinuesAfterPeerRejectsDelete(t *testing.T) {
resetTestGlobals()
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
defer cancel()
obj, disk, err := prepareFS(ctx)
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
if _, err := globalIAMSys.CreateUser(ctx, "heal-sync", req); err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.CreateUser(ctx, "heal-deleted", req); err != nil {
t.Fatal(err)
}
if err := globalIAMSys.DeleteUser(ctx, "heal-deleted", false); err != nil {
t.Fatal(err)
}
p, err := globalIAMSys.store.GetPolicy("readwrite")
if err != nil {
t.Fatal(err)
}
if _, err := globalIAMSys.SetPolicy(ctx, "heal-new-policy", p); err != nil {
t.Fatal(err)
}
var liveUpdates atomic.Int32
peer := func(id string, rejectDelete bool) *httptest.Server {
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
switch {
case strings.HasSuffix(r.URL.Path, "/metainfo"):
_ = json.NewEncoder(w).Encode(madmin.SRInfo{DeploymentID: id})
case r.URL.Path == "/minio/admin/v3/site-replication/peer/iam-revisions":
if r.Method == http.MethodGet {
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if rejectDelete && iamDeletionPath(item.SRIAMItem) != "" {
w.WriteHeader(http.StatusForbidden)
_, _ = w.Write([]byte(`{"Code":"AccessDenied","Message":"delete rejected"}`))
return
}
if id == "healthy" && item.Type == madmin.SRIAMItemPolicy && item.Name == "heal-new-policy" && len(item.Policy) > 0 {
liveUpdates.Add(1)
}
}
_ = json.NewEncoder(w).Encode(iamRevisionStatus{Version: iamRevisionProtocol, Node: "node-1", Instance: id, Digest: "fixture"})
default:
t.Errorf("unexpected peer request %s", r.URL.Path)
w.WriteHeader(http.StatusNotFound)
}
}))
}
healthy, rejected := peer("healthy", false), peer("rejected", true)
defer healthy.Close()
defer rejected.Close()
c := &SiteReplicationSys{enabled: true, state: srState{
ServiceAccountAccessKey: "heal-sync",
Peers: map[string]madmin.PeerInfo{
globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()},
"healthy": {Name: "healthy", DeploymentID: "healthy", Endpoint: healthy.URL},
"rejected": {Name: "rejected", DeploymentID: "rejected", Endpoint: rejected.URL},
},
}}
if err := c.healIAMSystem(ctx, obj); err == nil {
t.Fatal("deletion failure was not reported")
}
if liveUpdates.Load() == 0 {
t.Fatal("one peer rejecting a deletion blocked unrelated live IAM healing to a healthy peer")
}
}
func TestIAMRevocationLifecycle(t *testing.T) {
testIAMRevocationLifecycle(t, nil)
}
func TestIAMRevocationEtcdLifecycle(t *testing.T) {
endpoint := os.Getenv("SILO_TEST_IAM_REVOCATION_ETCD")
if endpoint == "" {
t.Skip("set SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint")
}
connection, err := etcd.New(etcd.Config{Endpoints: strings.Split(endpoint, ","), DialTimeout: 5 * time.Second})
if err != nil {
t.Fatal(err)
}
defer connection.Close()
// The facade borrows the connection's services. Close the owning client,
// not namespace.Watcher while IAM's canceled watch loop is winding down.
ctx, cancel := context.WithCancel(connection.Ctx())
defer cancel()
client := etcd.NewCtxClient(ctx, etcd.WithZapLogger(connection.GetLogger()))
prefix := fmt.Sprintf("/silo-revocation-test/%d/", time.Now().UnixNano())
client.KV = namespace.NewKV(connection.KV, prefix)
client.Watcher = namespace.NewWatcher(connection.Watcher, prefix)
client.Lease = connection.Lease
testIAMRevocationLifecycle(t, client)
}
func testIAMRevocationLifecycle(t *testing.T, client *etcd.Client) {
resetTestGlobals()
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
disks, err := getRandomDisks(1)
if err != nil {
t.Fatal(err)
}
disk := disks[0]
obj, _, err := initObjectLayer(ctx, mustGetPoolEndpoints(0, disks...))
if err == nil {
initAllSubsystems(ctx)
globalIAMSys.Init(ctx, obj, client, 2*time.Second)
}
if err != nil {
t.Fatal(err)
}
defer os.RemoveAll(disk)
defer obj.Shutdown(ctx)
defer resetTestGlobals()
sys, peer := globalIAMSys, &globalSiteReplicationSys
must := func(t *testing.T, err error) {
t.Helper()
if err != nil {
t.Fatal(err)
}
}
reload := func(t *testing.T) { t.Helper(); must(t, sys.store.LoadIAMCache(ctx, false)) }
req := madmin.AddOrUpdateUserReq{SecretKey: "valid-test-password", Status: madmin.AccountEnabled}
origin := UTCNow().Add(-time.Hour).Truncate(time.Millisecond)
createUser := func(t *testing.T, name string) {
t.Helper()
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, origin))
}
assertAbsent := func(t *testing.T, name string) {
t.Helper()
if _, ok := sys.GetUser(ctx, name); ok {
t.Fatalf("revoked credential %s is usable", name)
}
}
t.Run("replay after recreation", func(t *testing.T) {
testIAMRevocationReplayAfterRecreation(ctx, t, sys)
})
t.Run("origin timestamp and recreation", func(t *testing.T) {
name := "revocation-recreate"
createUser(t, name)
ui, ok := sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(origin) {
t.Fatalf("origin time changed: %v", ui.UpdatedAt)
}
must(t, sys.DeleteUser(ctx, name, false))
reload(t)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, time.Time{}))
assertAbsent(t, name)
newTime := UTCNow().Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, origin.Add(time.Second)))
reload(t)
ui, ok = sys.store.GetUser(name)
if !ok || !ui.UpdatedAt.Equal(newTime) || ui.RevokedBefore.IsZero() {
t.Fatalf("newer recreation lost, or deletion boundary missing: present=%v", ok)
}
})
t.Run("groups policies and mappings", func(t *testing.T) {
user, group, name := "revocation-member", "revocation-group", "revocation-policy"
createUser(t, user)
p, err := sys.store.GetPolicy("readwrite")
must(t, err)
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
add := &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
for _, isGroup := range []bool{false, true} {
entity := user
if isGroup {
entity = group
}
mp := &madmin.SRPolicyMapping{UserOrGroup: entity, Policy: name, UserType: int(regUser), IsGroup: isGroup}
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
_, err = sys.PolicyDBSet(ctx, entity, "", regUser, isGroup)
must(t, err)
must(t, peer.PeerPolicyMappingHandler(ctx, mp, origin))
if _, ok := sys.store.GetMappedPolicy(entity, isGroup); ok {
t.Fatal("old grant restored")
}
}
// This receiver never saw the member-removal event preceding deletion.
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, UTCNow()))
must(t, sys.DeletePolicy(ctx, name, true))
reload(t)
must(t, peer.PeerGroupInfoChangeHandler(ctx, add, origin))
must(t, peer.PeerAddPolicyHandler(ctx, name, &p, origin))
if _, err = sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("group restored: %v", err)
}
if _, err = sys.store.GetPolicyDoc(name); !errors.Is(err, errNoSuchPolicy) {
t.Fatalf("policy restored: %v", err)
}
paths, err := sys.store.listIAMConfigPaths(ctx)
must(t, err)
found := make(map[string]bool)
for _, path := range paths {
r, err := loadIAMRevision(ctx, sys.store, path)
must(t, err)
if item, ok := iamDeletionItem(path, r); ok {
found[iamDeletionPath(item)] = true
if item.UpdatedAt.IsZero() {
t.Fatal("undated delete replay")
}
}
}
for _, path := range []string{getGroupInfoPath(group), getPolicyDocPath(name), getMappedPolicyPath(user, regUser, false), getMappedPolicyPath(group, regUser, true)} {
if !found[path] {
t.Errorf("deletion missing from heal: %s", path)
}
}
})
t.Run("parent revokes service accounts and STS", func(t *testing.T) {
parent := "revocation-parent"
createUser(t, parent)
svc, svcAt, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "revocation-service", secretKey: "valid-service-password"})
must(t, err)
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
sts.ParentUser = parent
_, err = sys.SetTempUser(withIAMReplicationTime(ctx, origin), sts.AccessKey, sts, "readwrite")
must(t, err)
must(t, sys.DeleteUser(ctx, parent, false))
reload(t)
assertAbsent(t, parent)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// Recreate the parent, then deliver old child events from the offline site.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: svc.AccessKey, SecretKey: svc.SecretKey}}, svcAt))
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken, ParentPolicyMapping: "readwrite"}, origin))
reload(t)
assertAbsent(t, svc.AccessKey)
assertAbsent(t, sts.AccessKey)
// A freshly issued credential is still supported after deliberate recreation.
_, _, err = sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-service", secretKey: "valid-service-password"})
must(t, err)
if _, ok := sys.GetUser(ctx, "new-service"); !ok {
t.Fatal("fresh service account rejected")
}
})
t.Run("delete before first create", func(t *testing.T) {
name := "revocation-unseen"
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: name, IsDeleteReq: true}, UTCNow()))
createUser(t, name)
assertAbsent(t, name)
svc := "unseen-service"
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: svc}}, UTCNow()))
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "revocation-recreate", AccessKey: svc, SecretKey: "valid-service-password"}}, origin))
assertAbsent(t, svc)
})
t.Run("recreation arrives before revocation", func(t *testing.T) {
parent := "reordered-parent"
createUser(t, parent)
child, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: "reordered-child", secretKey: "valid-service-password"})
must(t, err)
newTime, deleteTime := UTCNow(), origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, newTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
assertAbsent(t, child.AccessKey)
reload(t)
assertAbsent(t, child.AccessKey)
u, ok := sys.GetUser(ctx, parent)
if !ok || !u.UpdatedAt.Equal(newTime) || !u.RevokedBefore.Equal(deleteTime) {
t.Fatal("reordered revocation damaged the new parent or lost its boundary")
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(parent, regUser))
must(t, err)
item, ok := iamDeletionItem(getUserIdentityPath(parent, regUser), r)
if !ok || !item.UpdatedAt.Equal(deleteTime) {
t.Fatal("recreation erased deletion replay")
}
})
t.Run("user cleanup does not supersede group deletion", func(t *testing.T) {
user, group := "cascade-user", "cascade-group"
createUser(t, user)
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, Members: []string{user}}}, origin))
// On the origin site the group was removed before the user, but the
// recovering receiver processes those independent events in reverse.
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: user, IsDeleteReq: true}, origin.Add(2*time.Minute)))
must(t, peer.PeerGroupInfoChangeHandler(ctx, &madmin.SRGroupInfo{UpdateReq: madmin.GroupAddRemove{Group: group, IsRemove: true}}, origin.Add(time.Minute)))
reload(t)
if _, err := sys.GetGroupDescription(group); !errors.Is(err, errNoSuchGroup) {
t.Fatalf("deleted group survived reordered cleanup: %v", err)
}
groups, err := sys.ListGroups(ctx)
must(t, err)
for _, name := range groups {
if name == group {
t.Fatal("deleted group listed")
}
}
})
t.Run("parent revocation covers later updates to existing children", func(t *testing.T) {
parent, key := "late-update-parent", "late-update-child"
createUser(t, parent)
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), parent, nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
must(t, err)
// This site missed the deletion and subsequently edited an old child.
_, err = sys.UpdateServiceAccount(withIAMReplicationTime(ctx, origin.Add(2*time.Minute)), key, updateServiceAccountOpts{description: "edited while the peer was offline"})
must(t, err)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, origin.Add(time.Minute)))
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(3*time.Minute)))
reload(t)
assertAbsent(t, key)
})
t.Run("old generation cannot return with a newer event timestamp", func(t *testing.T) {
parent := "generation-parent"
createUser(t, parent)
deleteTime := origin.Add(time.Minute)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, origin.Add(2*time.Minute)))
// Another offline site issued this child under the original parent,
// after this site's delete/recreate. Wall-clock ordering cannot identify it.
late := origin.Add(3 * time.Minute)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: "old-gen-service", SecretKey: "valid-service-password"}}, late))
secret, err := getTokenSigningKey()
must(t, err)
sts, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
must(t, peer.PeerSTSAccHandler(ctx, &madmin.SRSTSCredential{AccessKey: sts.AccessKey, SecretKey: sts.SecretKey, ParentUser: parent, SessionToken: sts.SessionToken}, late))
reload(t)
assertAbsent(t, "old-gen-service")
assertAbsent(t, sts.AccessKey)
// A local issuer knows the new boundary and signs it into both kinds
// of child. Untrusted inherited claims cannot select that boundary.
child, _, err := sys.NewServiceAccount(ctx, parent, nil, newServiceAccountOpts{accessKey: "new-gen-service", secretKey: "valid-service-password", claims: map[string]any{iamParentRevocationClaim: "forged"}})
must(t, err)
newClaims := map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}
must(t, setIAMParentRevocationClaim(ctx, sys.store, parent, newClaims))
fresh, err := auth.GetNewCredentialsWithMetadata(newClaims, secret)
must(t, err)
fresh.ParentUser = parent
_, err = sys.SetTempUser(ctx, fresh.AccessKey, fresh, "")
must(t, err)
reload(t)
// A periodic reload retains the STS cache. Explicitly clear it to
// exercise the cold credential load performed after process restart.
cache := sys.store.lock()
cache.iamSTSAccountsMap = make(map[string]UserIdentity)
sys.store.unlock()
for _, key := range []string{child.AccessKey, fresh.AccessKey} {
u, ok := sys.GetUser(ctx, key)
if !ok || !iamCredentialSurvivesRevocation(u.Credentials, deleteTime) {
t.Fatalf("new-generation credential %s rejected", key)
}
}
})
t.Run("late revocation preserves proven new-generation children", func(t *testing.T) {
for _, recreateFirst := range []bool{false, true} {
parent := fmt.Sprintf("gen-parent-%t", recreateFirst)
createUser(t, parent)
deleteTime, createTime := origin.Add(time.Minute), origin.Add(2*time.Minute)
if recreateFirst {
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
key := fmt.Sprintf("gen-child-%t", recreateFirst)
must(t, peer.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: parent, AccessKey: key, SecretKey: "valid-service-password", Claims: map[string]any{iamParentRevocationClaim: deleteTime.Format(time.RFC3339Nano)}}}, createTime.Add(time.Second)))
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, IsDeleteReq: true}, deleteTime))
if !recreateFirst {
assertAbsent(t, key)
must(t, peer.PeerIAMUserChangeHandler(ctx, &madmin.SRIAMUser{AccessKey: parent, UserReq: &req}, createTime))
}
reload(t)
if _, ok := sys.GetUser(ctx, key); !ok {
t.Fatal("late revocation deleted a child issued by the recreated parent")
}
}
})
t.Run("cold loading preserves site-signed STS", func(t *testing.T) {
parent := "cold-sts-parent"
createUser(t, parent)
secret := "site-signing-key-valid"
_, _, err := sys.NewServiceAccount(ctx, globalActiveCred.AccessKey, nil, newServiceAccountOpts{
accessKey: siteReplicatorSvcAcc, secretKey: secret, allowSiteReplicatorAccount: true,
})
must(t, err)
setReplication := func(enabled bool) {
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = enabled
globalSiteReplicationSys.Unlock()
globalSiteReplicatorCred.Set("")
}
setReplication(true)
defer setReplication(false)
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, secret)
must(t, err)
cred.ParentUser = parent
_, err = sys.SetTempUser(ctx, cred.AccessKey, cred, "")
must(t, err)
// IAM can load before the site replication manager during startup.
// A signing key that is not available yet must not delete live tokens.
setReplication(false)
for range 3 {
unverified := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, unverified)
if _, ok := unverified[cred.AccessKey]; ok {
t.Fatal("accepted STS before the signing key became available")
}
}
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath(cred.AccessKey, stsUser))
must(t, err)
if r.Credentials.SessionToken == "" {
t.Fatal("cold IAM load physically deleted a non-expired site-signed STS credential")
}
setReplication(true)
loaded := make(map[string]UserIdentity)
must(t, sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded))
if _, ok := loaded[cred.AccessKey]; !ok {
t.Fatal("STS credential did not recover when the signing key became available")
}
})
t.Run("unverifiable STS stay denied and expired STS are removed", func(t *testing.T) {
parent := "invalid-sts-parent"
createUser(t, parent)
// Keep the etcd watcher from cleaning half of the fixture before the
// second record is seeded; this subtest exercises the loader directly.
sys.store.lock()
defer sys.store.unlock()
for _, expired := range []bool{false, true} {
cred, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: parent}, "unavailable-test-signing-key")
must(t, err)
cred.ParentUser = parent
if expired {
cred.Expiration = UTCNow().Add(-time.Minute)
}
identityPath := getUserIdentityPath(cred.AccessKey, stsUser)
mappingPath := getMappedPolicyPath(cred.AccessKey, stsUser, false)
// Seed disk directly to exercise loading, including existing records
// whose key is unknown. The write API should not accept such tokens.
must(t, sys.store.saveIAMConfig(ctx, &UserIdentity{Version: 1, Credentials: cred, UpdatedAt: UTCNow()}, identityPath))
must(t, sys.store.saveIAMConfig(ctx, &MappedPolicy{Version: 1, Policies: "readwrite"}, mappingPath))
loaded := make(map[string]UserIdentity)
_ = sys.store.loadUser(ctx, cred.AccessKey, stsUser, loaded)
if _, ok := loaded[cred.AccessKey]; ok {
t.Fatalf("invalid STS accepted, expired=%t", expired)
}
for _, path := range []string{identityPath, mappingPath} {
var record map[string]any
err := sys.store.loadIAMConfig(ctx, &record, path)
if expired {
if !errors.Is(err, errConfigNotFound) {
t.Fatalf("expired STS data not cleaned up at %s: %v", path, err)
}
} else {
must(t, err)
}
}
}
})
}
+239
View File
@@ -0,0 +1,239 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio/internal/auth"
)
// The receiver missed a deletion and still has the previous service key.
// A later full snapshot must replace it, including when the owner changed.
func TestIAMServiceAccountRecreation(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
t.Run(backend, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
for _, parent := range []string{"old-owner", "new-owner"} {
_, err := sys.CreateUser(withIAMReplicationTime(ctx, origin), parent, madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
}
const key = "reusable-service"
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "old-owner", AccessKey: key, SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
oldIdentity, _ := sys.store.GetUser(key)
_, err := sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
mustIAM(t, err)
boundary, newer := origin.Add(time.Minute), origin.Add(2*time.Minute)
fresh := iamReplicationItem{SRIAMItem: madmin.SRIAMItem{Type: madmin.SRIAMItemSvcAcc, UpdatedAt: newer, SvcAccChange: &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "new-owner", AccessKey: key, SecretKey: "new-service-password", Status: auth.AccountOff}}}, RevokedBefore: boundary}
mustIAM(t, applyIAMReplicationItem(ctx, fresh))
// A sibling may have missed the notification of the committed
// replacement. An equal-version retry must refresh that cache too.
staleCache := sys.store.lock()
staleCache.iamUsersMap[key] = oldIdentity
sys.store.unlock()
mustIAM(t, applyIAMReplicationItem(ctx, fresh)) // duplicate delivery is acknowledged
if current, _ := sys.store.GetUser(key); current.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey {
t.Fatal("duplicate snapshot acknowledged without refreshing the stale cache")
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, &madmin.SRSvcAccChange{Delete: &madmin.SRSvcAccDelete{AccessKey: key}}, boundary))
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
u, ok := sys.store.GetUser(key)
if !ok || u.Credentials.SecretKey != fresh.SvcAccChange.Create.SecretKey || u.Credentials.ParentUser != "new-owner" || u.Credentials.Status != auth.AccountOff || !u.UpdatedAt.Equal(newer) || !u.RevokedBefore.Equal(boundary) {
t.Fatal("recreation did not retain the new identity, disabled status, source version and revocation")
}
if _, ok := sys.GetUser(ctx, key); ok {
t.Fatal("replicated disabled service can authenticate")
}
cache := sys.store.rlock()
_, mapped := cache.cachedMappedPolicy(key, svcUser, false)
sys.store.runlock()
if mapped {
t.Fatal("recreated service inherited an older mapping")
}
_, err = sys.PolicyDBSet(withIAMReplicationTime(ctx, origin), key, "readwrite", svcUser, false)
if !errors.Is(err, errIAMStaleUpdate) {
t.Fatalf("old service mapping replay was accepted: %v", err)
}
_, _, err = sys.NewServiceAccount(ctx, "new-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "local-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("local duplicate creation must remain rejected: %v", err)
}
// Outbound snapshots must carry the retained service boundary too.
out, err := globalSiteReplicationSys.replicationItem(ctx, fresh.SRIAMItem)
mustIAM(t, err)
if !out.RevokedBefore.Equal(boundary) {
t.Fatal("outbound service snapshot lost its revocation")
}
})
}
}
func TestIAMServiceAccountReplicationRejectsOtherCredentialKinds(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
_, err := sys.CreateUser(ctx, "builtin-collision", madmin.AddOrUpdateUserReq{SecretKey: "valid-user-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
secret, err := getTokenSigningKey()
mustIAM(t, err)
token, err := auth.GetNewCredentialsWithMetadata(map[string]any{"exp": UTCNow().Add(time.Hour).Unix(), parentClaim: "builtin-collision"}, secret)
mustIAM(t, err)
token.ParentUser = "builtin-collision"
_, err = sys.SetTempUser(ctx, token.AccessKey, token, "")
mustIAM(t, err)
for _, key := range []string{"builtin-collision", token.AccessKey} {
_, _, err := sys.NewServiceAccount(withIAMReplicationTime(ctx, UTCNow().Add(time.Minute)), "another-owner", nil, newServiceAccountOpts{accessKey: key, secretKey: "valid-service-password"})
if !errors.Is(err, errIAMServiceAccountNotAllowed) {
t.Fatalf("service replication replaced another credential kind: %v", err)
}
}
}
// SR configuration can be temporarily unreadable even though a service token
// is signed with its own valid secret. Do not acknowledge a failed cache load.
func TestIAMServiceAccountRetryReportsClaimLoadFailure(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t)
globalSiteReplicatorCred.RLock()
previousSigningKey := globalSiteReplicatorCred.secretKey
globalSiteReplicatorCred.RUnlock()
globalSiteReplicatorCred.Set("")
t.Cleanup(func() { globalSiteReplicatorCred.Set(previousSigningKey) })
_, err := sys.CreateUser(ctx, "retry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
opts := newServiceAccountOpts{accessKey: "retry-service", secretKey: "valid-service-password"}
_, _, err = sys.NewServiceAccount(ctx, "retry-owner", nil, opts)
mustIAM(t, err)
old, _ := sys.store.GetUser(opts.accessKey)
opts.secretKey = "replacement-service-password"
at, err := sys.UpdateServiceAccount(ctx, opts.accessKey, updateServiceAccountOpts{secretKey: opts.secretKey})
mustIAM(t, err)
cache := sys.store.lock()
cache.iamUsersMap[opts.accessKey] = old // Missed sibling notification.
sys.store.unlock()
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = true // No site-replicator credential is installed.
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
if err == nil {
t.Fatal("acknowledged service retry despite failed claims loading")
}
if _, ok := sys.store.GetUser(opts.accessKey); ok {
t.Fatal("failed cache refresh retained the superseded service secret")
}
globalSiteReplicationSys.Lock()
globalSiteReplicationSys.enabled = false
globalSiteReplicationSys.Unlock()
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, at), "retry-owner", nil, opts)
mustIAM(t, err)
}
// A delayed snapshot still has its original absolute expiration. Reapplying
// the local minimum issuance lifetime would leave the old unexpired key alive.
func TestIAMServiceAccountReplicationPreservesExpiration(t *testing.T) {
for _, backend := range []string{"object", "etcd"} {
for _, action := range []string{"create", "update"} {
for _, expired := range []bool{false, true} {
name := backend + "/" + action + "/near_expiry"
if expired {
name = backend + "/" + action + "/expired"
}
t.Run(name, func(t *testing.T) {
ctx, sys, _ := prepareIAMRevisionFixture(t, backend)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "expiry-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
old := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "old-service-password"}}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
expires := UTCNow().Add(time.Minute)
if expired {
expires = UTCNow().Add(-time.Minute)
}
change := &madmin.SRSvcAccChange{Create: &madmin.SRSvcAccCreate{Parent: "expiry-owner", AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
if action == "update" {
change = &madmin.SRSvcAccChange{Update: &madmin.SRSvcAccUpdate{AccessKey: "expiry-service", SecretKey: "new-service-password", Expiration: &expires}}
}
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, change, origin.Add(2*time.Minute)))
u, ok := sys.store.GetUser("expiry-service")
if !ok || u.Credentials.SecretKey != "new-service-password" || !u.Credentials.Expiration.Equal(expires) {
t.Fatal("delayed snapshot lost its new secret or absolute expiration")
}
_, allowed := sys.GetUser(ctx, "expiry-service")
if allowed == expired {
t.Fatal("credential validity disagrees with its absolute expiration")
}
mustIAM(t, sys.store.LoadIAMCache(ctx, false))
mustIAM(t, globalSiteReplicationSys.PeerSvcAccChangeHandler(ctx, old, origin))
r, err := loadIAMRevision(ctx, sys.store, getUserIdentityPath("expiry-service", svcUser))
mustIAM(t, err)
if r.Credentials.SecretKey == old.Create.SecretKey || (expired && !r.Deleted) {
t.Fatal("old non-expiring credential returned after reload")
}
_, _, err = sys.NewServiceAccount(ctx, "expiry-owner", nil, newServiceAccountOpts{accessKey: "local-expiry", secretKey: "valid-service-password", expiration: &expires})
if !errors.Is(err, errInvalidSvcAcctExpiration) {
t.Fatalf("local issuance lifetime check changed: %v", err)
}
})
}
}
}
}
// Status-only summaries intentionally omit secrets. Different revisions must
// still trigger live healing, and disabled identities must be eligible sources.
func TestIAMServiceAccountHealingNewerSnapshot(t *testing.T) {
ctx, sys, obj := prepareIAMRevisionFixture(t)
origin := UTCNow().Add(-time.Hour)
_, err := sys.CreateUser(ctx, "heal-svc-owner", madmin.AddOrUpdateUserReq{SecretKey: "valid-owner-password", Status: madmin.AccountEnabled})
mustIAM(t, err)
_, _, err = sys.NewServiceAccount(withIAMReplicationTime(ctx, origin), "heal-svc-owner", nil, newServiceAccountOpts{accessKey: "heal-service", secretKey: "valid-service-password"})
mustIAM(t, err)
at, err := sys.UpdateServiceAccount(ctx, "heal-service", updateServiceAccountOpts{status: auth.AccountOff})
mustIAM(t, err)
var sent atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/minio/health/live" {
return
}
if r.URL.Path != "/minio/admin/v3/site-replication/peer/iam-revisions" || r.Method != http.MethodPut {
t.Errorf("unexpected request %s %s", r.Method, r.URL.Path)
w.WriteHeader(http.StatusNotFound)
return
}
var batch iamRevisionBatch
if err := json.NewDecoder(r.Body).Decode(&batch); err != nil {
t.Error(err)
w.WriteHeader(http.StatusBadRequest)
return
}
for _, item := range batch.Items {
if item.SvcAccChange != nil && item.SvcAccChange.Create != nil && item.SvcAccChange.Create.AccessKey == "heal-service" && item.SvcAccChange.Create.Status == auth.AccountOff && item.UpdatedAt.Equal(at) {
sent.Add(1)
}
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "node", Instance: "boot", Digest: "fixture"}})
}))
defer server.Close()
peers := map[string]madmin.PeerInfo{globalDeploymentID(): {Name: "local", DeploymentID: globalDeploymentID()}, "remote": {Name: "remote", DeploymentID: "remote", Endpoint: server.URL}}
c := &SiteReplicationSys{enabled: true, state: srState{ServiceAccountAccessKey: "heal-svc-owner", Peers: peers}}
local := madmin.UserInfo{Status: madmin.AccountStatus(auth.AccountOff), UpdatedAt: at}
remote := local
remote.UpdatedAt = origin
if isUserInfoReplicated(2, 2, []madmin.UserInfo{local, remote}) {
t.Fatal("status-only summaries concealed different service revisions")
}
info := srStatusInfo{Sites: peers, UserStats: map[string]map[string]srUserStatsSummary{"heal-service": {globalDeploymentID(): {userInfo: srUserInfo{UserInfo: local}}, "remote": {SRUserStatsSummary: madmin.SRUserStatsSummary{UserInfoMismatch: true}, userInfo: srUserInfo{UserInfo: remote}}}}}
mustIAM(t, c.healUsers(ctx, obj, "heal-service", info))
if sent.Load() == 0 {
t.Fatal("disabled newer service snapshot was not healed")
}
}
+337 -221
View File
File diff suppressed because it is too large Load Diff
+56 -21
View File
@@ -605,10 +605,22 @@ func (sys *IAMSys) DeletePolicy(ctx context.Context, policyName string, notifyPe
return errServerNotInitialized
}
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
if err := checkConfig(ctx, globalObjectAPI, getPolicyDocPath(policyName)); err != nil && err == errConfigNotFound {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
if _, replicated := iamReplicationTime(ctx); !replicated && notifyPeers {
for _, v := range policy.DefaultPolicies {
if v.Name == policyName {
var err error
if objectStore, ok := sys.store.IAMStorageAPI.(*IAMObjectStore); ok {
err = checkConfig(ctx, objectStore.objAPI, getPolicyDocPath(policyName))
} else {
var r iamRevision
err = sys.store.loadIAMConfig(ctx, &r, getPolicyDocPath(policyName))
}
if errors.Is(err, errConfigNotFound) {
return fmt.Errorf("inbuilt policy `%s` not allowed to be deleted", policyName)
}
if err != nil {
return err
}
}
}
}
@@ -714,20 +726,30 @@ func (sys *IAMSys) DeleteUser(ctx context.Context, accessKey string, notifyPeers
return errServerNotInitialized
}
if err := sys.store.DeleteUser(ctx, accessKey, regUser); err != nil {
err := sys.store.DeleteUser(ctx, accessKey, regUser)
var cleanupErr *iamCommittedCleanupError
retained := errors.Is(err, errIAMRevocationRetained)
if errors.As(err, &cleanupErr) {
retained = cleanupErr.retained
} else if err != nil && !retained {
return err
}
// Notify all other MinIO peers to delete user.
// Publish the committed state even when dependent cleanup must be retried.
if notifyPeers && !sys.HasWatcher() {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
if retained {
sys.notifyForUser(ctx, accessKey, false)
} else {
for _, nerr := range globalNotificationSys.DeleteUser(ctx, accessKey) {
if nerr.Err != nil {
logger.GetReqInfo(ctx).SetTags("peerAddress", nerr.Host.String())
iamLogIf(ctx, nerr.Err)
}
}
}
}
if cleanupErr != nil {
return cleanupErr
}
return nil
}
@@ -1061,6 +1083,7 @@ type newServiceAccountOpts struct {
sessionPolicy *policy.Policy
accessKey string
secretKey string
status string // Used by replication snapshots; local creates default to enabled.
name, description string
expiration *time.Time
allowSiteReplicatorAccount bool // allow creating internal service account for site-replication.
@@ -1125,6 +1148,11 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
m[k] = v
}
}
if _, replicated := iamReplicationTime(ctx); !replicated {
if err := setIAMParentRevocationClaim(ctx, sys.store, parentUser, m); err != nil {
return auth.Credentials{}, time.Time{}, err
}
}
var accessKey, secretKey string
var err error
@@ -1143,12 +1171,19 @@ func (sys *IAMSys) NewServiceAccount(ctx context.Context, parentUser string, gro
cred.ParentUser = parentUser
cred.Groups = groups
cred.Status = string(auth.AccountOn)
switch opts.status {
case "", auth.AccountOn, string(madmin.AccountEnabled):
case auth.AccountOff, string(madmin.AccountDisabled):
cred.Status = auth.AccountOff
default:
return auth.Credentials{}, time.Time{}, errInvalidArgument
}
cred.Name = opts.name
cred.Description = opts.description
if opts.expiration != nil {
expirationInUTC := opts.expiration.UTC()
if err := validateSvcExpirationInUTC(expirationInUTC); err != nil {
if err := validateSvcExpirationInUTC(ctx, expirationInUTC); err != nil {
return auth.Credentials{}, time.Time{}, err
}
cred.Expiration = expirationInUTC
@@ -1379,7 +1414,7 @@ func (sys *IAMSys) DeleteServiceAccount(ctx context.Context, accessKey string, n
}
sa, ok := sys.store.GetUser(accessKey)
if !ok || !sa.Credentials.IsServiceAccount() {
if _, replicated := iamReplicationTime(ctx); (!ok || !sa.Credentials.IsServiceAccount()) && !replicated {
return nil
}
@@ -1483,8 +1518,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForExternalSSO(ctx context.Context) {
}
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// purgeExpiredCredentialsForLDAP - validates if local credentials are still
@@ -1512,8 +1547,8 @@ func (sys *IAMSys) purgeExpiredCredentialsForLDAP(ctx context.Context) {
return
}
// We ignore any errors
_ = sys.store.DeleteUsers(ctx, expiredUsers)
// Keep failed revocations visible so the next purge can retry.
iamLogIf(ctx, sys.store.DeleteUsers(ctx, expiredUsers))
}
// updateGroupMembershipsForLDAP - updates the list of groups associated with the credential.
@@ -1934,12 +1969,12 @@ func (sys *IAMSys) RemoveUsersFromGroup(ctx context.Context, group string, membe
}
updatedAt, err = sys.store.RemoveUsersFromGroup(ctx, group, members)
if err != nil {
var cleanupErr *iamCommittedCleanupError
if err != nil && !errors.As(err, &cleanupErr) {
return updatedAt, err
}
sys.notifyForGroup(ctx, group)
return updatedAt, nil
return updatedAt, err
}
// SetGroupStatus - enable/disabled a group
+292
View File
@@ -0,0 +1,292 @@
// Copyright (c) 2026 mr javad seydi
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"errors"
"fmt"
"slices"
"testing"
"time"
)
func multipartUploadKeys(uploads []MultipartInfo) []string {
keys := make([]string, len(uploads))
for i := range uploads {
keys[i] = uploads[i].Object
}
return keys
}
func requireMultipartUploadKeys(t *testing.T, got ListMultipartsInfo, want ...string) {
t.Helper()
if keys := multipartUploadKeys(got.Uploads); !slices.Equal(keys, want) {
t.Fatalf("uploads = %v, want %v", keys, want)
}
}
func TestListMultipartUploadsS3Compatibility(t *testing.T) {
obj, dirs, err := prepareErasureSets32(t.Context())
if err != nil {
t.Fatal(err)
}
z := obj.(*erasureServerPools)
setMultipartListingTestMode(t, false)
t.Cleanup(func() {
z.Shutdown(t.Context())
removeRoots(dirs)
})
const bucket = "multipart-list-compat"
if err = z.MakeBucket(t.Context(), bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
objects := []string{"t/a_b/p1", "t/a_b/p2", "t/c_d/p1", "u/x"}
sets := z.serverPools[0]
firstSet := sets.getHashedSetIndex(objects[0])
if !slices.ContainsFunc(objects[1:], func(object string) bool {
return sets.getHashedSetIndex(object) != firstSet
}) {
for n := 0; ; n++ {
object := fmt.Sprintf("v/cross-set-%d", n)
if sets.getHashedSetIndex(object) != firstSet {
objects = append(objects, object)
break
}
}
}
uploadIDs := make(map[string]string, len(objects))
for _, object := range objects {
mp, err := z.NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatalf("NewMultipartUpload(%q): %v", object, err)
}
uploadIDs[object] = mp.UploadID
}
// Durable multipart metadata, rather than this node-local cache, must be
// authoritative after a restart or when another node handles the request.
z.mpCache.Range(func(uploadID string, _ MultipartInfo) bool {
z.mpCache.Delete(uploadID)
return true
})
all, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, all, objects...)
if all.IsTruncated || all.NextKeyMarker != "" || all.NextUploadIDMarker != "" {
t.Fatalf("complete listing has truncation state: %+v", all)
}
first, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 1)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, first, objects[0])
if !first.IsTruncated || first.NextKeyMarker != objects[0] || first.NextUploadIDMarker != uploadIDs[objects[0]] {
t.Fatalf("first page markers = (%q, %q, %t), want (%q, %q, true)",
first.NextKeyMarker, first.NextUploadIDMarker, first.IsTruncated,
objects[0], uploadIDs[objects[0]])
}
rest, err := z.ListMultipartUploads(t.Context(), bucket, "", first.NextKeyMarker, first.NextUploadIDMarker, "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, rest, objects[1:]...)
afterKey, err := z.ListMultipartUploads(t.Context(), bucket, "", objects[1], "", "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, afterKey, objects[2:]...)
prefixed, err := z.ListMultipartUploads(t.Context(), bucket, "t/", "", "", "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, prefixed, objects[:3]...)
nested, err := z.ListMultipartUploads(t.Context(), bucket, "t/a_b/", "", "", "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, nested, objects[:2]...)
grouped, err := z.ListMultipartUploads(t.Context(), bucket, "t/", "", "", SlashSeparator, 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, grouped)
if want := []string{"t/a_b/", "t/c_d/"}; !slices.Equal(grouped.CommonPrefixes, want) {
t.Fatalf("common prefixes = %v, want %v", grouped.CommonPrefixes, want)
}
groupPage, err := z.ListMultipartUploads(t.Context(), bucket, "t/", "", "", SlashSeparator, 1)
if err != nil {
t.Fatal(err)
}
if want := []string{"t/a_b/"}; !slices.Equal(groupPage.CommonPrefixes, want) {
t.Fatalf("first common-prefix page = %v, want %v", groupPage.CommonPrefixes, want)
}
if !groupPage.IsTruncated || groupPage.NextKeyMarker != "t/a_b/" || groupPage.NextUploadIDMarker != "" {
t.Fatalf("common-prefix page markers = (%q, %q, %t)",
groupPage.NextKeyMarker, groupPage.NextUploadIDMarker, groupPage.IsTruncated)
}
groupRest, err := z.ListMultipartUploads(t.Context(), bucket, "t/", groupPage.NextKeyMarker, "", SlashSeparator, 1)
if err != nil {
t.Fatal(err)
}
if want := []string{"t/c_d/"}; !slices.Equal(groupRest.CommonPrefixes, want) {
t.Fatalf("second common-prefix page = %v, want %v", groupRest.CommonPrefixes, want)
}
if groupRest.IsTruncated {
t.Fatalf("last common-prefix page is truncated: %+v", groupRest)
}
part, err := z.PutObjectPart(t.Context(), bucket, objects[0], uploadIDs[objects[0]], 1,
mustGetPutObjReader(t, bytes.NewBufferString("part"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
completed, err := z.CompleteMultipartUpload(t.Context(), bucket, objects[0], uploadIDs[objects[0]],
[]CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
for _, key := range []string{multipartMetaBucket, multipartMetaObject} {
if _, ok := completed.UserDefined[key]; ok {
t.Errorf("completed object retained upload-only metadata %q", key)
}
}
// Simulate an upload written by a pre-upgrade server. Its key cannot be
// recovered by scanning the hashed namespace. Strict listing must fail;
// the old path is available only through an explicit migration setting.
legacyObject := objects[1]
er := sets.getHashedSet(legacyObject)
fi, metadata, err := er.checkUploadIDExists(t.Context(), bucket, legacyObject, uploadIDs[legacyObject], true)
if err != nil {
t.Fatal(err)
}
for i := range metadata {
delete(metadata[i].Metadata, multipartMetaBucket)
delete(metadata[i].Metadata, multipartMetaObject)
}
if _, err = writeAllMetadata(t.Context(), er.getDisks(), bucket, minioMetaMultipartBucket,
er.getUploadIDDir(bucket, legacyObject, uploadIDs[legacyObject]), metadata, fi.WriteQuorum(er.defaultWQuorum())); err != nil {
t.Fatal(err)
}
_, err = z.ListMultipartUploads(t.Context(), bucket, legacyObject, "", "", "", 100)
if !errors.Is(err, errMultipartListingLegacy) {
t.Fatalf("legacy strict listing: %v", err)
}
setMultipartListingTestMode(t, true)
legacy, err := z.ListMultipartUploads(t.Context(), bucket, legacyObject, "", "", "", 100)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, legacy, legacyObject)
}
func TestPaginateMultipartUploads(t *testing.T) {
base := time.Unix(100, 0)
id1, id2, id3 := multipartListingTestID(base, 1), multipartListingTestID(base.Add(time.Second), 2), multipartListingTestID(base, 3)
uploads := []MultipartInfo{
{Bucket: "bucket", Object: "b", UploadID: id3, Initiated: base},
{Bucket: "bucket", Object: "a", UploadID: id2, Initiated: base.Add(time.Second)},
{Bucket: "bucket", Object: "a", UploadID: id1, Initiated: base},
{Bucket: "bucket", Object: "a", UploadID: id1, Initiated: base}, // duplicate discovery
}
first := paginateMultipartUploads(uploads, "", "", "", "", 1)
requireMultipartUploadKeys(t, first, "a")
if !first.IsTruncated || first.NextKeyMarker != "a" || first.NextUploadIDMarker != id1 {
t.Fatalf("first page = %+v", first)
}
second := paginateMultipartUploads(uploads, "", first.NextKeyMarker, first.NextUploadIDMarker, "", 1)
if len(second.Uploads) != 1 || second.Uploads[0].Object != "a" || second.Uploads[0].UploadID != id2 {
t.Fatalf("second page uploads = %+v", second.Uploads)
}
if !second.IsTruncated || second.NextKeyMarker != "a" || second.NextUploadIDMarker != id2 {
t.Fatalf("second page = %+v", second)
}
last := paginateMultipartUploads(uploads, "", second.NextKeyMarker, second.NextUploadIDMarker, "", 1)
requireMultipartUploadKeys(t, last, "b")
if last.IsTruncated || last.NextKeyMarker != "" || last.NextUploadIDMarker != "" {
t.Fatalf("last page = %+v", last)
}
missingUploadMarker := paginateMultipartUploads(uploads, "", "a", multipartListingTestID(base.Add(2*time.Second), 4), "", 10)
requireMultipartUploadKeys(t, missingUploadMarker, "b")
if err := checkListMultipartArgs(t.Context(), "bucket", "", "", "not-base64=", ""); err != nil {
t.Fatalf("upload-id-marker without key-marker must be ignored: %v", err)
}
overLimit := make([]MultipartInfo, maxUploadsList+1)
for i := range overLimit {
overLimit[i] = MultipartInfo{Bucket: "bucket", Object: fmt.Sprintf("%04d", i), UploadID: fmt.Sprint(i)}
}
capped := paginateMultipartUploads(overLimit, "", "", "", "", maxUploadsList+1)
if capped.MaxUploads != maxUploadsList || len(capped.Uploads) != maxUploadsList || !capped.IsTruncated {
t.Fatalf("over-limit page = MaxUploads %d, uploads %d, truncated %t",
capped.MaxUploads, len(capped.Uploads), capped.IsTruncated)
}
}
func TestListMultipartUploadsGlobalPageAcrossPools(t *testing.T) {
z, bucket := consistencyPools(t)
setMultipartListingTestMode(t, false)
objects := []string{"a/one", "b/two", "c/three", "d/four"}
for i, object := range objects {
if _, err := z.serverPools[i%len(z.serverPools)].NewMultipartUpload(t.Context(), bucket, object, ObjectOptions{}); err != nil {
t.Fatalf("NewMultipartUpload(%q): %v", object, err)
}
}
z.mpCache.Range(func(uploadID string, _ MultipartInfo) bool {
z.mpCache.Delete(uploadID)
return true
})
page, err := z.ListMultipartUploads(t.Context(), bucket, "", "", "", "", 2)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, page, objects[:2]...)
if !page.IsTruncated || page.NextKeyMarker != objects[1] {
t.Fatalf("first global page = %+v", page)
}
rest, err := z.ListMultipartUploads(t.Context(), bucket, "", page.NextKeyMarker, page.NextUploadIDMarker, "", 2)
if err != nil {
t.Fatal(err)
}
requireMultipartUploadKeys(t, rest, objects[2:]...)
if rest.IsTruncated {
t.Fatalf("last global page is truncated: %+v", rest)
}
}
+14
View File
@@ -34,6 +34,10 @@ const (
sinceLastSyncMillis = "since_last_sync_millis"
syncFailures = "sync_failures"
syncSuccesses = "sync_successes"
revocationRecords = "revocation_records"
revocationHealFailures = "revocation_heal_failures"
revocationHealDurationMillis = "revocation_heal_duration_millis"
revocationHealLastSuccess = "revocation_heal_last_success_timestamp_seconds"
)
var (
@@ -47,10 +51,20 @@ var (
sinceLastSyncMillisMD = NewCounterMD(sinceLastSyncMillis, "Time (in milliseconds) since last successful IAM data sync.")
syncFailuresMD = NewCounterMD(syncFailures, "Number of failed IAM data syncs since server start.")
syncSuccessesMD = NewCounterMD(syncSuccesses, "Number of successful IAM data syncs since server start.")
revocationRecordsMD = NewGaugeMD(revocationRecords, "Retained IAM deletion records and revocation boundaries in this node's index.")
revocationHealFailuresMD = NewCounterMD(revocationHealFailures, "Failed IAM revocation convergence passes since server start.")
revocationHealDurationMillisMD = NewGaugeMD(revocationHealDurationMillis, "Duration of the last IAM revocation convergence pass in milliseconds.")
revocationHealLastSuccessMD = NewGaugeMD(revocationHealLastSuccess, "Unix timestamp of the last successful IAM revocation convergence pass.")
)
// loadClusterIAMMetrics - `MetricsLoaderFn` for cluster IAM metrics.
func loadClusterIAMMetrics(_ context.Context, m MetricValues, _ *metricsCache) error {
if globalIAMSys.Initialized() {
m.Set(revocationRecords, float64(globalIAMSys.store.revisionIndex().count()))
}
m.Set(revocationHealFailures, float64(globalSiteReplicationSys.iamRevisionMetrics.healFailures.Load()))
m.Set(revocationHealDurationMillis, float64(globalSiteReplicationSys.iamRevisionMetrics.healDurationMillis.Load()))
m.Set(revocationHealLastSuccess, float64(globalSiteReplicationSys.iamRevisionMetrics.healLastSuccess.Load()))
m.Set(lastSyncDurationMillis, float64(atomic.LoadUint64(&globalIAMSys.LastRefreshDurationMilliseconds)))
pluginAuthNMetrics := globalAuthNPlugin.Metrics()
m.Set(pluginAuthnServiceFailedRequestsMinute, float64(pluginAuthNMetrics.FailedRequests))
+4
View File
@@ -323,6 +323,10 @@ func newMetricGroups(r *prometheus.Registry) *metricsV3Collection {
sinceLastSyncMillisMD,
syncFailuresMD,
syncSuccessesMD,
revocationRecordsMD,
revocationHealFailuresMD,
revocationHealDurationMillisMD,
revocationHealLastSuccessMD,
},
loadClusterIAMMetrics,
)
+6 -1
View File
@@ -20,6 +20,7 @@ package cmd
import (
"context"
"encoding/base64"
"errors"
"runtime"
"strings"
@@ -83,7 +84,8 @@ func checkListMultipartArgs(ctx context.Context, bucket, prefix, keyMarker, uplo
if err := checkListObjsArgs(ctx, bucket, prefix, keyMarker); err != nil {
return err
}
if uploadIDMarker != "" {
// S3 ignores upload-id-marker when key-marker is absent.
if uploadIDMarker != "" && keyMarker != "" {
if HasSuffix(keyMarker, SlashSeparator) {
return InvalidUploadIDKeyCombination{
UploadIDMarker: uploadIDMarker,
@@ -96,6 +98,9 @@ func checkListMultipartArgs(ctx context.Context, bucket, prefix, keyMarker, uplo
UploadID: uploadIDMarker,
}
}
if _, ok := multipartMarkerTime(uploadIDMarker); !ok {
return InvalidArgument{Bucket: bucket, Object: keyMarker, Err: errors.New("upload-id-marker must contain a native multipart upload ID")}
}
}
return nil
}
+113
View File
@@ -0,0 +1,113 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"encoding/base64"
"net/http"
"reflect"
"strconv"
"strings"
"testing"
"time"
xhttp "github.com/minio/minio/internal/http"
)
func TestPutOptsFromHeadersReplicationTimestamps(t *testing.T) {
stamp := time.Date(2026, 9, 15, 1, 2, 3, 123456789, time.UTC)
context := base64.StdEncoding.EncodeToString([]byte(`{"purpose":"tag-replication"}`))
for _, encryption := range []struct {
name string
headers map[string]string
}{
{name: "none"},
{name: "SSE-S3", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionAES}},
{name: "SSE-KMS", headers: map[string]string{xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS}},
{name: "SSE-KMS-context", headers: map[string]string{
xhttp.AmzServerSideEncryption: xhttp.AmzEncryptionKMS, xhttp.AmzServerSideEncryptionKmsID: "tag-replication-key",
xhttp.AmzServerSideEncryptionKmsContext: context,
}},
{name: "SSE-C", headers: ssecKeyHeaders([]byte("01234567890123456789012345678901"), false)},
} {
t.Run(encryption.name, func(t *testing.T) {
for _, trusted := range []bool{false, true} {
t.Run("trusted="+strconv.FormatBool(trusted), func(t *testing.T) {
for _, tagging := range []struct {
name, header string
want time.Time
invalid bool
}{
{name: "absent"},
{name: "nanoseconds", header: stamp.Format(time.RFC3339Nano), want: stamp},
{name: "offset-whitespace", header: " " + stamp.In(time.FixedZone("UTC+8", 8*60*60)).Format(time.RFC3339Nano) + " ", want: stamp},
{name: "invalid", header: "not-a-timestamp", invalid: true},
} {
t.Run(tagging.name, func(t *testing.T) {
for _, metadata := range []map[string]string{nil, {"x-amz-meta-test": "kept"}} {
hdr := make(http.Header)
wantEncryption := make(http.Header)
for key, value := range encryption.headers {
hdr.Set(key, value)
wantEncryption.Set(key, value)
}
hdr.Set(xhttp.MinIOSourceTaggingTimestamp, tagging.header)
hdr.Set(xhttp.MinIOSourceMTime, stamp.Add(-time.Hour).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectRetentionTimestamp, stamp.Add(-time.Minute).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceObjectLegalHoldTimestamp, stamp.Add(-time.Second).Format(time.RFC3339Nano))
hdr.Set(xhttp.MinIOSourceETag, "source-etag")
opts, err := putOptsFromHeaders(t.Context(), hdr, metadata, trusted)
if trusted && tagging.invalid {
if err == nil || !strings.Contains(err.Error(), xhttp.MinIOSourceTaggingTimestamp) {
t.Fatalf("malformed trusted timestamp: got %v", err)
}
continue
}
if err != nil {
t.Fatal(err)
}
wantTag, wantMTime, wantRetention, wantLegalhold, wantETag := time.Time{}, time.Time{}, time.Time{}, time.Time{}, ""
if trusted {
wantTag, wantMTime = tagging.want, stamp.Add(-time.Hour)
wantRetention, wantLegalhold, wantETag = stamp.Add(-time.Minute), stamp.Add(-time.Second), "source-etag"
}
if !opts.ReplicationSourceTaggingTimestamp.Equal(wantTag) {
t.Errorf("tag timestamp=%s, want %s", opts.ReplicationSourceTaggingTimestamp, wantTag)
}
if !opts.MTime.Equal(wantMTime) || !opts.ReplicationSourceRetentionTimestamp.Equal(wantRetention) ||
!opts.ReplicationSourceLegalholdTimestamp.Equal(wantLegalhold) || opts.PreserveETag != wantETag || opts.ReplicationRequest != trusted {
t.Error("other source fields did not preserve the replication trust boundary")
}
if opts.UserDefined == nil || (metadata != nil && !reflect.DeepEqual(opts.UserDefined, metadata)) {
t.Errorf("metadata=%v, want nonnil map preserving %v", opts.UserDefined, metadata)
}
gotEncryption := make(http.Header)
if opts.ServerSideEncryption != nil {
opts.ServerSideEncryption.Marshal(gotEncryption)
}
if !reflect.DeepEqual(gotEncryption, wantEncryption) {
t.Errorf("SSE headers=%v, want %v", gotEncryption, wantEncryption)
}
}
})
}
})
}
})
}
}
+3 -3
View File
@@ -452,11 +452,11 @@ func putOptsFromHeaders(ctx context.Context, hdr http.Header, metadata map[strin
MTime: mtime,
PreserveETag: etag,
ReplicationRequest: trustedReplication,
// The Object Lock timestamps order replicated retention and legal
// hold updates. Dropping them here would leave every update on an
// SSE-KMS destination unordered.
// These timestamps order replicated retention, legal hold and tagging
// updates on an SSE-KMS destination.
ReplicationSourceLegalholdTimestamp: lholdtimestmp,
ReplicationSourceRetentionTimestamp: retaintimestmp,
ReplicationSourceTaggingTimestamp: taggingtimestmp,
}
return op, nil
}
+143
View File
@@ -0,0 +1,143 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/crypto"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
// TestAPICopyObjectReplicaTaggingTimestampUnderKMS covers signed replica COPY
// requests through encryption, metadata replacement and disk persistence. Both
// the single-disk and 16-disk fixtures are single-pool backends.
func TestAPICopyObjectReplicaTaggingTimestampUnderKMS(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPICopyObjectReplicaTaggingTimestampUnderKMS})
}
func testAPICopyObjectReplicaTaggingTimestampUnderKMS(obj ObjectLayer, instance, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
// Ignore the host free-space percentage while retaining real disk I/O.
for _, pool := range obj.(*erasureServerPools).serverPools {
for _, set := range pool.sets {
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
defer func() { set.getDisks = original }()
}
}
oldKMS, oldAuto := GlobalKMS, globalAutoEncryption
GlobalKMS = kms.NewStub("replica-tags-key")
globalAutoEncryption = false
defer func() { GlobalKMS, globalAutoEncryption = oldKMS, oldAuto }()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
const tsKey = ReservedMetadataPrefixLower + TaggingTimestamp
stamp := time.Date(2026, 9, 15, 1, 0, 0, 123456789, time.UTC)
for _, mode := range []string{"none", "explicit-sse-s3", "explicit-kms", "auto-kms", "bucket-kms"} {
t.Run(instance+"/"+mode, func(t *testing.T) {
globalAutoEncryption = mode == "auto-kms"
if mode == "bucket-kms" {
sseXML := []byte(`<ServerSideEncryptionConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><Rule><ApplyServerSideEncryptionByDefault><SSEAlgorithm>aws:kms</SSEAlgorithm><KMSMasterKeyID>replica-tags-key</KMSMasterKeyID></ApplyServerSideEncryptionByDefault></Rule></ServerSideEncryptionConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketSSEConfig, sseXML); err != nil {
t.Fatal(err)
}
}
const data = "encrypted replica copy remains readable"
oi, err := obj.PutObject(t.Context(), bucket, mode, mustGetPutObjReader(t, bytes.NewReader([]byte(data)), int64(len(data)), "", ""), ObjectOptions{
Versioned: true, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=old", tsKey: stamp.Format(time.RFC3339Nano)},
})
if err != nil {
t.Fatal(err)
}
for _, event := range []struct {
name, tags, wantTags string
delta, wantDelta time.Duration
missingTimestamp bool
}{
{"newer", "key=new", "key=new", 2, 2, false},
{"stale", "key=stale", "key=new", 1, 2, false},
{"duplicate", "key=new", "key=new", 2, 2, false},
{"newer-again", "key=latest", "key=latest", 3, 3, false},
{"missing-timestamp", "key=unordered", "key=latest", 0, 3, true},
} {
headers := map[string]string{
xhttp.AmzCopySource: "/" + bucket + "/" + mode + "?versionId=" + oi.VersionID,
xhttp.AmzMetadataDirective: "REPLACE", xhttp.AmzTagDirective: "REPLACE",
xhttp.AmzObjectTagging: event.tags, xhttp.MinIOSourceReplicationRequest: "true",
xhttp.AmzBucketReplicationStatus: "REPLICA", xhttp.MinIOSourceTaggingTimestamp: stamp.Add(event.delta).Format(time.RFC3339Nano),
xhttp.MinIOSourceMTime: oi.ModTime.Format(time.RFC3339Nano), xhttp.MinIOSourceETag: oi.ETag,
}
if event.missingTimestamp {
delete(headers, xhttp.MinIOSourceTaggingTimestamp)
}
if mode == "explicit-sse-s3" {
headers[xhttp.AmzServerSideEncryption] = xhttp.AmzEncryptionAES
}
if mode == "explicit-kms" {
headers[xhttp.AmzServerSideEncryption] = "aws:kms"
headers[xhttp.AmzServerSideEncryptionKmsID] = "replica-tags-key"
}
req, err := newTestSignedRequestV4(http.MethodPut, "/"+bucket+"/"+mode+"?versionId="+oi.VersionID, 0, nil, creds.AccessKey, creds.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusOK {
t.Fatalf("%s: COPY %d %s", event.name, w.Code, w.Body.String())
}
got, err := obj.GetObjectInfo(t.Context(), bucket, mode, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
t.Logf("%s: tags=%q timestamp=%q kms=%v", event.name, got.UserTags, got.UserDefined[tsKey], crypto.S3KMS.IsEncrypted(got.UserDefined))
if got.UserTags != event.wantTags || got.UserDefined[tsKey] != stamp.Add(event.wantDelta).Format(time.RFC3339Nano) {
t.Errorf("%s: incorrect persisted tags/timestamp", event.name)
}
wantKMS := mode == "explicit-kms" || mode == "auto-kms" || mode == "bucket-kms"
if crypto.S3KMS.IsEncrypted(got.UserDefined) != wantKMS || crypto.S3.IsEncrypted(got.UserDefined) != (mode == "explicit-sse-s3") {
t.Errorf("%s: unexpected destination encryption", event.name)
}
if got.VersionID != oi.VersionID {
t.Errorf("%s: version=%q, want %q", event.name, got.VersionID, oi.VersionID)
}
req, err = newTestSignedRequestV4(http.MethodGet, "/"+bucket+"/"+mode+"?versionId="+oi.VersionID, 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
w = httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusOK || w.Body.String() != data {
t.Fatalf("%s: GET %d %q", event.name, w.Code, w.Body.String())
}
}
})
}
}
+4 -1
View File
@@ -240,7 +240,10 @@ func checkPreconditionsPUT(ctx context.Context, w http.ResponseWriter, r *http.R
// updated. The predicate is the incoming request's restored SSE-C metadata,
// not what the destination happens to hold.
ssecReplica := isReplicaTrusted(r.Context()) && crypto.SSEC.IsEncrypted(opts.UserDefined)
if etagMatch && vidMatch && !ssecReplica {
// Matching content does not imply that its tag revision was delivered.
// Keep client preconditions above; relax only the internal duplicate check.
newerTags := isReplicaTrusted(r.Context()) && olderThan(objInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp], opts.ReplicationSourceTaggingTimestamp)
if etagMatch && vidMatch && !ssecReplica && !newerTags {
writeHeaders()
writeErrorResponse(ctx, w, errorCodes.ToAPIErr(ErrPreconditionFailed), r.URL)
return true
+30 -19
View File
@@ -1797,6 +1797,7 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
// source timestamp is newer than the stored one, and a stale update must
// leave the stored state in place instead of erasing it.
storedLock := storedObjectLockState(srcInfo.UserDefined)
storedTagTimestamp := srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
srcInfo.UserDefined, err = getCpObjMetadataFromHeader(ctx, r, srcInfo.UserDefined, allowReplicationMetadata)
if err != nil {
@@ -1814,23 +1815,29 @@ func (api objectAPIHandlers) CopyObjectHandler(w http.ResponseWriter, r *http.Re
}
}
if objTags != "" {
lastTaggingTimestamp := srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp]
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceTaggingTimestamp
if !srcTimestamp.IsZero() {
ondiskTimestamp, err := time.Parse(time.RFC3339Nano, lastTaggingTimestamp)
// update tagging metadata only if replica timestamp is newer than what's on disk
if err != nil || (err == nil && !ondiskTimestamp.After(srcTimestamp)) {
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
}
}
} else {
if dstOpts.ReplicationRequest {
srcTimestamp := dstOpts.ReplicationSourceTaggingTimestamp
if !srcTimestamp.IsZero() {
// An empty value with a timestamp is an ordered deletion. Recheck
// the captured state even if metadata REPLACE rebuilt the map.
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = srcTimestamp.UTC().Format(time.RFC3339Nano)
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
reconcileStoredObjectTags(srcInfo.UserDefined, srcInfo.UserTags, storedTagTimestamp)
} else {
srcInfo.UserDefined[xhttp.AmzObjectTagging] = srcInfo.UserTags
if storedTagTimestamp != "" {
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = storedTagTimestamp
} else {
delete(srcInfo.UserDefined, ReservedMetadataPrefixLower+TaggingTimestamp)
}
}
} else {
srcInfo.UserDefined[xhttp.AmzObjectTagging] = objTags
srcInfo.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
// SSE-C rotation snapshots reserved metadata before the tag decision. Its
// later merge must not put the old timestamp back over the accepted state.
delete(encMetadata, ReservedMetadataPrefixLower+TaggingTimestamp)
srcInfo.UserDefined = filterReplicationStatusMetadata(srcInfo.UserDefined)
srcInfo.UserDefined = objectlock.FilterObjectLockMetadata(srcInfo.UserDefined, true, true)
@@ -2313,6 +2320,9 @@ func (api objectAPIHandlers) PutObjectHandler(w http.ResponseWriter, r *http.Req
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
if opts.ReplicationRequest && !opts.ReplicationSourceTaggingTimestamp.IsZero() {
metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = opts.ReplicationSourceTaggingTimestamp.UTC().Format(time.RFC3339Nano)
}
actualSize := size
var idxCb func() []byte
@@ -3760,11 +3770,11 @@ func (api objectAPIHandlers) PutObjectTaggingHandler(w http.ResponseWriter, r *h
}
dsc := mustReplicate(ctx, bucket, object, getMustReplicateOptions(objInfo.UserDefined, tagsStr, objInfo.ReplicationStatus, replication.MetadataReplicationType, opts))
stamp := UTCNow().Format(time.RFC3339Nano)
opts.UserDefined = map[string]string{ReservedMetadataPrefixLower + TaggingTimestamp: stamp}
if dsc.ReplicateAny() {
opts.UserDefined = make(map[string]string)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = UTCNow().Format(time.RFC3339Nano)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = stamp
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationStatus] = dsc.PendingStatus()
opts.UserDefined[ReservedMetadataPrefixLower+TaggingTimestamp] = UTCNow().Format(time.RFC3339Nano)
}
// Put object tags
@@ -3863,9 +3873,10 @@ func (api objectAPIHandlers) DeleteObjectTaggingHandler(w http.ResponseWriter, r
}
dsc := mustReplicate(ctx, bucket, object, oi.getMustReplicateOptions(replication.MetadataReplicationType, opts))
stamp := UTCNow().Format(time.RFC3339Nano)
opts.UserDefined = map[string]string{ReservedMetadataPrefixLower + TaggingTimestamp: stamp}
if dsc.ReplicateAny() {
opts.UserDefined = make(map[string]string)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = UTCNow().Format(time.RFC3339Nano)
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] = stamp
opts.UserDefined[ReservedMetadataPrefixLower+ReplicationStatus] = dsc.PendingStatus()
}
+4
View File
@@ -312,6 +312,10 @@ func (api objectAPIHandlers) NewMultipartUploadHandler(w http.ResponseWriter, r
writeErrorResponse(ctx, w, toAPIError(ctx, err), r.URL)
return
}
// Completion orders the upload's persisted tag state under the object lock.
if opts.ReplicationRequest && !opts.ReplicationSourceTaggingTimestamp.IsZero() {
metadata[ReservedMetadataPrefixLower+TaggingTimestamp] = opts.ReplicationSourceTaggingTimestamp.UTC().Format(time.RFC3339Nano)
}
if r.Header.Get(xhttp.IfMatch) != "" {
opts.HasIfMatch = true
+6 -2
View File
@@ -204,7 +204,9 @@ func (s *peerRESTServer) DeleteServiceAccountHandler(mss *grid.MSS) (np grid.NoP
return np, grid.NewRemoteErr(errors.New("service account name is missing"))
}
if err := globalIAMSys.DeleteServiceAccount(context.Background(), accessKey, false); err != nil {
ctx, cancel := context.WithTimeout(GlobalContext, defaultContextTimeout)
defer cancel()
if err := globalIAMSys.LoadServiceAccount(ctx, accessKey); err != nil {
return np, grid.NewRemoteErr(err)
}
@@ -274,7 +276,9 @@ func (s *peerRESTServer) LoadUserHandler(mss *grid.MSS) (np grid.NoPayload, nerr
userType = stsUser
}
if err = globalIAMSys.LoadUser(context.Background(), objAPI, accessKey, userType); err != nil {
ctx, cancel := context.WithTimeout(GlobalContext, defaultContextTimeout)
defer cancel()
if err = globalIAMSys.LoadUser(ctx, objAPI, accessKey, userType); err != nil {
return np, grid.NewRemoteErr(err)
}
+276
View File
@@ -0,0 +1,276 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"archive/tar"
"bytes"
"compress/gzip"
"encoding/xml"
"maps"
"net/http"
"net/http/httptest"
"net/url"
"reflect"
"strings"
"testing"
"github.com/minio/minio/internal/auth"
xhttp "github.com/minio/minio/internal/http"
)
// Exercise authenticated handlers and actual disk metadata, including the
// response headers consumers see after replication has completed.
func TestAPIReplicaContentEncoding(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: testAPIReplicaContentEncoding})
}
func testAPIReplicaContentEncoding(obj ObjectLayer, instance, bucket string, router http.Handler, owner auth.Credentials, t *testing.T) {
ordinary := newObjectAttributesAuthzUser(t, instance, bucket, `"s3:PutObject","s3:GetObject"`)
replicator := newObjectAttributesAuthzUser(t, instance, bucket, `"s3:PutObject","s3:GetObject","s3:ReplicateObject"`)
for _, mode := range []string{"ordinary", "untrusted-marker", "replica"} {
for _, tc := range []struct{ name, wire, want string }{
{"bare", "aws-chunked", ""}, {"mixed", "aws-chunked,gzip", "gzip"}, {"gzip", "gzip", "gzip"},
} {
for _, operation := range []string{"put", "copy-replace", "multipart"} {
t.Run(instance+"/"+mode+"/"+tc.name+"/"+operation, func(t *testing.T) {
object := mode + "/" + tc.name + "/" + operation
payload := replicaEncodingPayload(t, tc.want)
creds := ordinary
headers := map[string]string{xhttp.ContentEncoding: tc.wire, xhttp.ContentType: "application/octet-stream", "X-Amz-Meta-Source": "encoding-test"}
if mode != "ordinary" {
headers[xhttp.MinIOSourceReplicationRequest] = "true"
}
if mode == "replica" {
creds = replicator
headers[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
send := func(method, target string, data []byte, hdrs map[string]string) *httptest.ResponseRecorder {
t.Helper()
req, err := newTestSignedRequestV4(method, target, int64(len(data)), bytes.NewReader(data), creds.AccessKey, creds.SecretKey, hdrs)
if err != nil {
t.Fatal(err)
}
return replicaEncodingServe(t, router, req, http.StatusOK)
}
switch operation {
case "put":
if strings.Contains(tc.wire, "aws-chunked") {
req := replicaEncodingStream(t, getPutObjectURL("", bucket, object), payload, creds, headers)
replicaEncodingServe(t, router, req, http.StatusOK)
} else {
send(http.MethodPut, getPutObjectURL("", bucket, object), payload, headers)
}
case "copy-replace":
source := object + "-source"
if _, err := obj.PutObject(t.Context(), bucket, source, mustGetPutObjReader(t, bytes.NewReader(payload), int64(len(payload)), "", ""), ObjectOptions{}); err != nil {
t.Fatal(err)
}
headers[xhttp.AmzCopySource] = url.QueryEscape("/" + bucket + "/" + source)
headers[xhttp.AmzMetadataDirective] = replaceDirective
send(http.MethodPut, getCopyObjectURL("", bucket, object), nil, headers)
case "multipart":
rec := send(http.MethodPost, getNewMultipartURL("", bucket, object), nil, headers)
var init InitiateMultipartUploadResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &init); err != nil {
t.Fatal(err)
}
// Part/completion metadata must not replace the encoding saved at initiation.
partHeaders := map[string]string{xhttp.ContentEncoding: "br"}
if mode == "replica" {
partHeaders[xhttp.MinIOSourceReplicationRequest] = "true"
partHeaders[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
part := send(http.MethodPut, getPutObjectPartURL("", bucket, object, init.UploadID, "1"), payload, partHeaders)
partETags := part.Header()[xhttp.ETag]
if len(partETags) != 1 {
t.Fatalf("missing part ETag: %#v", part.Header())
}
complete, err := xml.Marshal(CompleteMultipartUpload{Parts: []CompletePart{{PartNumber: 1, ETag: canonicalizeETag(partETags[0])}}})
if err != nil {
t.Fatal(err)
}
send(http.MethodPost, getCompleteMultipartUploadURL("", bucket, object, init.UploadID), complete, partHeaders)
}
assertReplicaEncodingObject(t, obj, router, owner, bucket, object, tc.want, payload)
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if got := info.UserDefined[xhttp.AmzBucketReplicationStatus]; (got == "REPLICA") != (mode == "replica") {
t.Errorf("replica status %q for mode %s", got, mode)
}
if info.ContentType != "application/octet-stream" {
t.Errorf("content-type=%q", info.ContentType)
}
if value, ok := caseInsensitiveMap(info.UserDefined).Lookup("x-amz-meta-source"); !ok || value != "encoding-test" {
t.Errorf("user metadata lost: %#v", info.UserDefined)
}
})
}
}
}
t.Run(instance+"/unauthorized-replica", func(t *testing.T) {
object := "denied-replica"
req := replicaEncodingStream(t, getPutObjectURL("", bucket, object), []byte("denied"), ordinary, map[string]string{xhttp.ContentEncoding: "aws-chunked", xhttp.MinIOSourceReplicationRequest: "true", xhttp.AmzBucketReplicationStatus: "REPLICA"})
rec := replicaEncodingServe(t, router, req, http.StatusForbidden)
var response APIErrorResponse
if err := xml.Unmarshal(rec.Body.Bytes(), &response); err != nil {
t.Fatal(err)
}
if response.Code != "AccessDenied" {
t.Fatalf("expected permission denial, got %s", response.Code)
}
if _, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{}); err == nil {
t.Error("denied replica created an object")
}
})
}
func replicaEncodingPayload(t *testing.T, encoding string) []byte {
t.Helper()
data := bytes.Repeat([]byte("replica encoding payload\n"), 128)
if encoding != "gzip" {
return data
}
var b bytes.Buffer
w := gzip.NewWriter(&b)
if _, err := w.Write(data); err != nil {
t.Fatal(err)
}
if err := w.Close(); err != nil {
t.Fatal(err)
}
return b.Bytes()
}
func replicaEncodingStream(t *testing.T, target string, data []byte, creds auth.Credentials, headers map[string]string) *http.Request {
t.Helper()
const chunkSize = 64
body := bytes.NewReader(data)
req, err := newTestStreamingRequest(http.MethodPut, target, int64(len(data)), chunkSize, body)
if err != nil {
t.Fatal(err)
}
for k, v := range headers {
req.Header.Set(k, v)
}
now := UTCNow()
signature, err := signStreamingRequest(req, creds.AccessKey, creds.SecretKey, now)
if err != nil {
t.Fatal(err)
}
req, err = assembleStreamingChunks(req, body, chunkSize, creds.SecretKey, signature, now)
if err != nil {
t.Fatal(err)
}
return req
}
func replicaEncodingServe(t *testing.T, router http.Handler, req *http.Request, want int) *httptest.ResponseRecorder {
t.Helper()
rec := httptest.NewRecorder()
router.ServeHTTP(rec, req)
if rec.Code != want {
t.Fatalf("%s %s: status=%d want=%d body=%s", req.Method, req.URL, rec.Code, want, rec.Body.String())
}
return rec
}
func assertReplicaEncodingObject(t *testing.T, obj ObjectLayer, router http.Handler, creds auth.Credentials, bucket, object, encoding string, data []byte) {
t.Helper()
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if info.ContentEncoding != encoding {
t.Errorf("persisted content-encoding=%q want=%q", info.ContentEncoding, encoding)
}
if encoding == "" {
if _, present := info.UserDefined["content-encoding"]; present {
t.Error("transport-only content-encoding key persisted")
}
}
for _, method := range []string{http.MethodGet, http.MethodHead} {
req, err := newTestSignedRequestV4(method, getPutObjectURL("", bucket, object), 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
rec := replicaEncodingServe(t, router, req, http.StatusOK)
if got := rec.Header().Get(xhttp.ContentEncoding); got != encoding {
t.Errorf("%s content-encoding=%q want=%q", method, got, encoding)
}
if encoding == "" {
if _, present := rec.Header()[xhttp.ContentEncoding]; present {
t.Errorf("%s sent an empty/transport encoding header", method)
}
}
if method == http.MethodGet && !bytes.Equal(rec.Body.Bytes(), data) {
t.Errorf("GET body differs: got %d bytes want %d", rec.Body.Len(), len(data))
}
}
}
func TestAPISnowballReplicaContentEncoding(t *testing.T) {
defer DetectTestLeak(t)()
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
for _, tc := range []struct {
name string
pax map[string]string
want string
}{
{name: "no-pax"},
{name: "pax-without-encoding", pax: map[string]string{"minio.metadata.Content-Type": "application/octet-stream"}},
{name: "pax-bare", pax: map[string]string{"minio.metadata.Content-Encoding": "aws-chunked"}},
{name: "pax-mixed", pax: map[string]string{"minio.metadata.Content-Encoding": "aws-chunked,gzip"}, want: "gzip"},
} {
t.Run(instance+"/"+tc.name, func(t *testing.T) {
object := "snowball/" + tc.name
data := replicaEncodingPayload(t, tc.want)
var archive bytes.Buffer
tw := tar.NewWriter(&archive)
if err := tw.WriteHeader(&tar.Header{Name: object, Mode: 0o600, Size: int64(len(data)), PAXRecords: tc.pax}); err != nil {
t.Fatal(err)
}
if _, err := tw.Write(data); err != nil {
t.Fatal(err)
}
if err := tw.Close(); err != nil {
t.Fatal(err)
}
var ordinaryMetadata map[string]string
// An unauthorized entry in a REPLICA request is rejected. Compare the
// same archive across ordinary and authorized replica requests instead.
for _, replica := range []bool{false, true} {
headers := map[string]string{
xhttp.ContentEncoding: "aws-chunked", xhttp.AmzSnowballExtract: "true",
xhttp.ContentType: "application/x-tar", xhttp.CacheControl: "max-age=123",
"X-Amz-Meta-Archive": "outer-request",
}
if replica {
headers[xhttp.MinIOSourceReplicationRequest] = "true"
headers[xhttp.AmzBucketReplicationStatus] = "REPLICA"
}
req := replicaEncodingStream(t, getPutObjectURL("", bucket, "archive.tar"), archive.Bytes(), creds, headers)
replicaEncodingServe(t, router, req, http.StatusOK)
assertReplicaEncodingObject(t, obj, router, creds, bucket, object, tc.want, data)
info, err := obj.GetObjectInfo(t.Context(), bucket, object, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
metadata := maps.Clone(info.UserDefined)
for _, key := range []string{xhttp.AmzBucketReplicationStatus, ReservedMetadataPrefixLower + ReplicaStatus, ReservedMetadataPrefixLower + ReplicaTimestamp, "etag"} {
delete(metadata, key)
}
if !replica {
ordinaryMetadata = metadata
} else if !reflect.DeepEqual(metadata, ordinaryMetadata) {
t.Errorf("replica inherited ordinary archive metadata: got %#v want %#v", metadata, ordinaryMetadata)
}
}
})
}
}})
}
+5 -19
View File
@@ -106,6 +106,7 @@ func TestReplicateDeleteMarkerTargetSemantics(t *testing.T) {
func testReplicateDeleteMarkerPurge(obj ObjectLayer, instanceType, bucket string, router http.Handler, creds auth.Credentials, t *testing.T, legacy bool) {
ctx := t.Context()
defer replicationTestCapacity(obj)()
const arn = "arn:minio:replication::af470089-d354-4473-934c-9e1f52f6da89:bucket"
const name = "marker"
version := mustGetUUID()
@@ -208,27 +209,12 @@ func testReplicateDeleteMarkerPurge(obj ObjectLayer, instanceType, bucket string
t.Errorf("purge scheduled as marker creation: version=%q marker=%q", deletion.VersionID, deletion.DeleteMarkerVersionID)
}
if legacy {
// Reproduce the old producer's state and let the existing scanner/heal
// path recover it. Upgrades must also finish purges already left pending.
// Old task shapes must complete directly; scanner/MRF recovery is
// covered separately by TestReplicationMRFMarkerRecovery.
deletion.VersionID, deletion.DeleteMarkerVersionID = "", version
replicateDelete(ctx, deletion, obj)
oi, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version, Versioned: true})
if oi.VersionPurgeStatus != replication.VersionPurgePending {
t.Fatalf("legacy source purge = %s, want PENDING", oi.VersionPurgeStatus)
}
targets, err := globalBucketTargetSys.ListBucketTargets(ctx, bucket)
if err != nil {
t.Fatal(err)
}
queueReplicationHeal(ctx, bucket, oi, replicationConfig{Config: &cfg, remotes: targets}, 0)
select {
case op := <-worker:
deletion = op.(DeletedObjectReplicationInfo)
case <-time.After(time.Second):
t.Fatal("legacy pending purge was not scheduled for healing")
}
}
result := replicateDelete(context.Background(), deletion, obj)
result := replicateDelete(context.Background(), deletion, markerPurgeUpdateLayer{ObjectLayer: obj, t: t})
if result.VersionPurgeStatus() != replication.VersionPurgeComplete {
t.Errorf("remote purge result = %s, want COMPLETE", result.VersionPurgeStatus())
}
+628
View File
@@ -0,0 +1,628 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/logger"
loghttp "github.com/minio/minio/internal/logger/target/http"
"github.com/minio/minio/internal/once"
xnet "github.com/pgsty/silo-pkg/v3/net"
)
type markerRecoveryCase struct {
name string
legacy, creation, lostReply, partial, exhaust, lockFirst, invalidMRF, replicaSource, offline, unrecordedPurge bool
targets int
}
func TestReplicationMRFMarkerRecovery(t *testing.T) {
for _, tc := range []markerRecoveryCase{
{name: "canonical", targets: 1},
{name: "lock-failure", lockFirst: true, targets: 1},
{name: "invalid-MRF-metadata", invalidMRF: true, targets: 1},
{name: "legacy", legacy: true, targets: 1},
{name: "unrecorded-purge", unrecordedPurge: true, targets: 2},
{name: "two-targets", legacy: true, targets: 2},
{name: "two-targets-offline", offline: true, targets: 2},
{name: "replica-source", replicaSource: true, targets: 1},
{name: "lost-reply-canonical", lostReply: true, targets: 1},
{name: "lost-reply-legacy", legacy: true, lostReply: true, targets: 1},
{name: "creation", creation: true, targets: 1},
{name: "partial-creation-block", partial: true, targets: 2},
{name: "retry-budget-and-scanner", exhaust: true, targets: 1},
} {
t.Run(tc.name, func(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, endpoints: []string{"DeleteObject"}, objAPITest: func(obj ObjectLayer, backend, bucket string, router http.Handler, creds auth.Credentials, t *testing.T) {
testReplicationMRFMarkerRecovery(t, obj, backend, bucket, router, creds, tc)
}})
})
}
}
type markerRecoveryTarget struct {
arn, bucket string
client *minio.Client
reject, loseReply atomic.Bool
deletes atomic.Int32
}
// Only adapt host capacity accounting, using the existing real-disk adapter.
// Reads, object metadata, MRF files and writes all still use the fixture disks.
func replicationTestCapacity(obj ObjectLayer) func() {
var restore []func()
for _, pool := range obj.(*erasureServerPools).serverPools {
for _, set := range pool.sets {
original := set.getDisks
disks := append([]StorageAPI(nil), original()...)
for i, disk := range disks {
if disk != nil {
disks[i] = tagTestCapacityDisk{StorageAPI: disk}
}
}
set.getDisks = func() []StorageAPI { return disks }
restore = append(restore, func() { set.getDisks = original })
}
}
return func() {
for _, fn := range restore {
fn()
}
}
}
func testReplicationMRFMarkerRecovery(t *testing.T, obj ObjectLayer, backend, bucket string, router http.Handler, creds auth.Credentials, tc markerRecoveryCase) {
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
defer replicationTestCapacity(obj)()
stats := NewReplicationStats(ctx, nil)
oldStats := globalReplicationStats.Swap(stats)
defer globalReplicationStats.Store(oldStats)
oldPool := globalReplicationPool
defer func() { globalReplicationPool = oldPool }()
const name = "marker"
version := mustGetUUID()
creationTime := UTCNow().Add(-time.Hour).Truncate(time.Second)
if _, err := globalBucketMetadataSys.Update(ctx, bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
if _, err := obj.PutObject(ctx, bucket, name, mustGetPutObjReader(t, bytes.NewReader([]byte("data")), 4, "", ""), ObjectOptions{Versioned: true}); err != nil {
t.Fatal(err)
}
var targets []*markerRecoveryTarget
cfg := replication.Config{}
creationStates := make(map[string]replication.StatusType)
var creationInternal string
for i := 0; i < tc.targets; i++ {
target := &markerRecoveryTarget{arn: "arn:minio:replication::" + mustGetUUID() + ":bucket", bucket: getRandomBucketName()}
if err := obj.MakeBucket(ctx, target.bucket, MakeBucketOptions{}); err != nil {
t.Fatal(err)
}
if _, err := obj.PutObject(ctx, target.bucket, name, mustGetPutObjReader(t, bytes.NewReader([]byte("data")), 4, "", ""), ObjectOptions{Versioned: true}); err != nil {
t.Fatal(err)
}
if !tc.creation {
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, ReplicationRequest: true, MTime: creationTime}
opts.SetReplicaStatus(replication.Replica)
if _, err := obj.DeleteObject(ctx, target.bucket, name, opts); err != nil {
t.Fatal(err)
}
}
remote := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
opts := ObjectOptions{VersionID: r.URL.Query().Get("versionId"), Versioned: true}
switch r.Method {
case http.MethodHead:
oi, err := obj.GetObjectInfo(r.Context(), target.bucket, name, opts)
if oi.DeleteMarker {
w.Header().Set(xhttp.AmzDeleteMarker, "true")
w.Header().Set(xhttp.AmzVersionID, oi.VersionID)
}
if err != nil {
writeErrorResponseHeadersOnly(w, toAPIError(r.Context(), err))
return
}
w.WriteHeader(http.StatusOK)
case http.MethodDelete:
target.deletes.Add(1)
if got := r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"; got != tc.creation {
t.Errorf("purge/creation wire flag=%v creation=%v", got, tc.creation)
}
if opts.VersionID == "" {
t.Error("missing remote versionId")
}
if target.reject.Load() {
w.WriteHeader(http.StatusForbidden)
fmt.Fprint(w, `<Error><Code>AccessDenied</Code></Error>`)
return
}
opts.DeleteMarker = r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"
opts.ReplicationRequest = true
opts.SetReplicaStatus(replication.Replica)
_, err := obj.DeleteObject(r.Context(), target.bucket, name, opts)
if err != nil && !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
writeErrorResponse(r.Context(), w, toAPIError(r.Context(), err), r.URL)
return
}
if target.loseReply.Swap(false) {
conn, _, err := w.(http.Hijacker).Hijack()
if err != nil {
t.Error(err)
return
}
conn.Close()
return
}
w.WriteHeader(http.StatusNoContent)
default:
t.Errorf("unexpected remote method %s", r.Method)
}
}))
defer remote.Close()
client, err := minio.New(strings.TrimPrefix(remote.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
target.client = client
globalBucketTargetSys.Lock()
globalBucketTargetSys.arnRemotesMap[target.arn] = arnTarget{Client: &TargetClient{Client: client, ARN: target.arn, Bucket: target.bucket}, lastRefresh: UTCNow()}
globalBucketTargetSys.targetsMap[bucket] = append(globalBucketTargetSys.targetsMap[bucket], madmin.BucketTarget{Arn: target.arn, TargetBucket: target.bucket})
globalBucketTargetSys.Unlock()
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[client.EndpointURL().Host] = epHealth{Online: true}
globalBucketTargetSys.hMutex.Unlock()
rule := configs[0].Rules[0]
rule.Priority = i + 1
rule.Destination = replication.Destination{ARN: target.arn, Bucket: target.bucket}
cfg.Rules = append(cfg.Rules, rule)
creationStates[target.arn] = replication.Completed
creationInternal += target.arn + "=COMPLETED;"
targets = append(targets, target)
}
if tc.targets == 1 {
cfg.RoleArn = targets[0].arn
}
if !tc.creation {
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, MTime: creationTime, DeleteReplication: ReplicationState{Targets: creationStates, ReplicationStatusInternal: creationInternal, ReplicationTimeStamp: creationTime}}
if tc.replicaSource {
opts.DeleteReplication = ReplicationState{}
opts.SetReplicaStatus(replication.Replica)
}
if _, err := obj.DeleteObject(ctx, bucket, name, opts); err != nil {
t.Fatal(err)
}
}
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatal(err)
}
meta.replicationConfig = &cfg
globalBucketMetadataSys.Set(bucket, meta)
newPool := func() *ReplicationPool {
p := &ReplicationPool{ctx: ctx, objLayer: obj, workers: []chan ReplicationWorkerOperation{make(chan ReplicationWorkerOperation, 8)}, stats: stats, mrfSaveCh: make(chan MRFReplicateEntry, 8)}
globalReplicationPool = once.NewSingleton[ReplicationPool]()
globalReplicationPool.Set(p)
return p
}
p := newPool()
receive := func() DeletedObjectReplicationInfo {
select {
case op := <-p.workers[0]:
d, ok := op.(DeletedObjectReplicationInfo)
if !ok {
t.Fatalf("wrong queued operation %T", op)
}
return d
case <-time.After(3 * time.Second):
t.Fatal("no marker task from actual MRF/handler/scanner queue")
return DeletedObjectReplicationInfo{}
}
}
uri := "/" + bucket + "/" + name
if !tc.creation {
uri += "?versionId=" + version
}
req, err := newTestSignedRequestV4(http.MethodDelete, uri, 0, nil, creds.AccessKey, creds.SecretKey, nil)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, req)
if w.Code != http.StatusNoContent {
t.Fatalf("source DELETE status %d: %s", w.Code, w.Body)
}
deletion := receive()
if tc.creation {
version = deletion.DeleteMarkerVersionID
} else if deletion.VersionID != version || deletion.DeleteMarkerVersionID != "" {
t.Fatalf("current producer emitted noncanonical purge: %+v", deletion)
}
if tc.unrecordedPurge {
// Exercise a task carrying purge state while the disk marker still has
// only creation metadata. Recreate only this isolated fixture marker.
if _, err := obj.DeleteObject(ctx, bucket, name, ObjectOptions{VersionID: version, Versioned: true}); err != nil {
t.Fatal(err)
}
opts := ObjectOptions{VersionID: version, Versioned: true, DeleteMarker: true, MTime: creationTime, DeleteReplication: ReplicationState{Targets: creationStates, ReplicationStatusInternal: creationInternal, ReplicationTimeStamp: creationTime}}
if _, err := obj.DeleteObject(ctx, bucket, name, opts); err != nil {
t.Fatal(err)
}
}
if tc.legacy {
deletion.VersionID, deletion.DeleteMarkerVersionID = "", version
}
if tc.partial {
deletion.TargetArn = targets[len(targets)-1].arn
}
if tc.invalidMRF {
testReplicationMRFInvalidLookups(ctx, t, obj, p, newPool, deletion)
return
}
var auditStatuses <-chan string
if tc.name == "canonical" {
var stopAudit func()
auditStatuses, stopAudit = replicationTestAudit(ctx, t, bucket)
defer stopAudit()
}
assertAudit := func(want string) {
if auditStatuses == nil {
return
}
select {
case got := <-auditStatuses:
if got != want {
t.Fatalf("replication audit status=%q want %q", got, want)
}
case <-time.After(3 * time.Second):
t.Fatal("missing replication audit event")
}
}
deletion.OpType = replication.HealReplicationType // Exercise operation-status statistics too.
failing := targets[len(targets)-1]
failing.reject.Store(!tc.lostReply)
failing.loseReply.Store(tc.lostReply)
if tc.offline {
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[failing.client.EndpointURL().Host] = epHealth{Online: false}
globalBucketTargetSys.hMutex.Unlock()
}
before, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
beforeCreation := before.ReplicationStatusInternal
beforeStamp := before.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp]
beforeReplica := before.UserDefined[ReservedMetadataPrefixLower+ReplicaStatus]
beforeReplicaStamp := before.UserDefined[ReservedMetadataPrefixLower+ReplicaTimestamp]
if !tc.creation && !tc.replicaSource && (len(replicationStatusesMap(beforeCreation)) != tc.targets || beforeStamp == "") {
t.Fatalf("missing seeded creation block: %+v", before)
}
if tc.replicaSource && (beforeReplica != "REPLICA" || beforeReplicaStamp == "") {
t.Fatalf("missing replica block: %+v", before)
}
updateObj := obj
if !tc.creation {
updateObj = markerPurgeUpdateLayer{ObjectLayer: obj, t: t}
}
rounds := 2
if tc.lostReply {
rounds = 1
}
if tc.exhaust {
rounds = mrfRetryLimit + 1
}
for round := 1; round <= rounds; round++ {
callObj := updateObj
if tc.lockFirst && round == 1 {
callObj = markerLockFailureLayer{ObjectLayer: updateObj}
}
result := replicateDelete(ctx, deletion, callObj)
switch {
case tc.lockFirst && round == 1:
if len(result.Targets) != 0 {
t.Fatal("lock failure attempted a target")
}
case tc.creation:
if result.ReplicationStatus() != replication.Failed {
t.Fatalf("creation result: %+v", result)
}
case result.VersionPurgeStatus() != replication.VersionPurgeFailed:
t.Fatalf("purge failure result: %+v", result)
}
assertAudit("FAILED")
if round == 1 && !tc.lockFirst && !tc.creation {
stats.RLock()
failed := stats.Cache[bucket].Stats[failing.arn].FailStats.SinceUptime
stats.RUnlock()
if failed.Count != 1 || failed.Bytes != 0 {
t.Fatalf("purge failure stats=%+v, want count 1 and zero bytes", failed)
}
}
oi, err := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
if !isErrMethodNotAllowed(err) || !oi.DeleteMarker {
t.Fatalf("source marker metadata/405 missing: %+v %v", oi, err)
}
if !tc.creation && (oi.ReplicationStatusInternal != beforeCreation || oi.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp] != beforeStamp) {
t.Fatalf("purge rewrote creation block: before=%q/%q after=%q/%q", beforeCreation, beforeStamp, oi.ReplicationStatusInternal, oi.UserDefined[ReservedMetadataPrefixLower+ReplicationTimestamp])
}
if !tc.creation && (oi.UserDefined[ReservedMetadataPrefixLower+ReplicaStatus] != beforeReplica || oi.UserDefined[ReservedMetadataPrefixLower+ReplicaTimestamp] != beforeReplicaStamp) {
t.Fatal("purge rewrote replica block")
}
if tc.targets == 2 && !tc.partial {
purges := versionPurgeStatusesMap(oi.VersionPurgeStatusInternal)
if purges[targets[0].arn] != replication.VersionPurgeComplete || purges[failing.arn] != replication.VersionPurgeFailed {
t.Fatalf("incorrect persisted target states: %v", purges)
}
if targets[0].deletes.Load() != 1 {
t.Fatalf("successful target resent %d times", targets[0].deletes.Load())
}
}
if tc.partial {
select {
case <-p.mrfSaveCh:
default:
t.Fatal("partial failure missing MRF")
}
dsc := deletion.ReplicationState.ReplicateDecisionStr
deletion.ReplicationState = oi.ReplicationState()
deletion.ReplicationState.ReplicateDecisionStr = dsc
continue
}
if tc.exhaust && round > mrfRetryLimit {
if len(p.mrfSaveCh) != 0 || atomic.LoadUint64(&stats.mrfStats.TotalDroppedCount) != 1 {
t.Fatalf("retry budget not applied: queued=%d drops=%d", len(p.mrfSaveCh), stats.mrfStats.TotalDroppedCount)
}
break
}
var entry MRFReplicateEntry
select {
case entry = <-p.mrfSaveCh:
default:
t.Fatal("failure did not enter MRF")
}
if entry.RetryCount != round || entry.versionID != version {
t.Fatalf("MRF entry lost identity/budget: %+v, round=%d", entry, round)
}
p.saveMRFEntries(ctx, map[string]MRFReplicateEntry{entry.versionID: entry})
record, err := p.loadMRF()
if err != nil || len(record.Entries) != 1 {
t.Fatalf("disk MRF missing: %+v %v", record, err)
}
if got := record.Entries[version]; got.RetryCount != round || got.Object != name || got.Bucket != bucket {
t.Fatalf("disk MRF mismatch: %+v", got)
}
p.saveMRFEntries(ctx, record.Entries) // loadMRF consumes the file; each replay uses a fresh disk record.
p = newPool() // no in-memory entries carried to the replacement pool.
if err := p.queueMRFHeal(); err != nil {
t.Fatal(err)
}
deletion = receive()
if deletion.RetryCount != round {
t.Fatalf("MRF retry count=%d want %d", deletion.RetryCount, round)
}
if !tc.creation && (deletion.VersionID != version || deletion.DeleteMarkerVersionID != "") {
t.Fatalf("MRF emitted wrong purge: %+v", deletion)
}
}
if tc.partial {
return
}
failing.reject.Store(false)
globalBucketTargetSys.hMutex.Lock()
globalBucketTargetSys.hc[failing.client.EndpointURL().Host] = epHealth{Online: true}
globalBucketTargetSys.hMutex.Unlock()
if tc.exhaust {
oi, _ := obj.GetObjectInfo(ctx, bucket, name, ObjectOptions{VersionID: version})
QueueReplicationHeal(ctx, bucket, oi, 0) // Separately prove the existing scanner fallback after MRF exhaustion.
deletion = receive()
if deletion.RetryCount != 0 {
t.Fatal("scanner did not start a fresh retry budget")
}
}
result := replicateDelete(ctx, deletion, updateObj)
assertAudit("COMPLETED")
if tc.creation {
if result.ReplicationStatus() != replication.Completed {
t.Fatalf("creation recovery failed: %+v", result)
}
} else if result.VersionPurgeStatus() != replication.VersionPurgeComplete {
t.Fatalf("purge recovery failed: %+v", result)
}
for _, b := range append([]string{bucket}, func() []string {
var b []string
for _, target := range targets {
b = append(b, target.bucket)
}
return b
}()...) {
oi, err := obj.GetObjectInfo(ctx, b, name, ObjectOptions{VersionID: version})
if tc.creation {
if !oi.DeleteMarker || !isErrMethodNotAllowed(err) {
t.Fatalf("creation missing in %s: %+v %v", b, oi, err)
}
} else if !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
t.Fatalf("purge left marker in %s: %+v %v", b, oi, err)
}
}
if !tc.creation {
// Re-deliver an ambiguous purge directly. An absent marker must stay absent.
duplicate := deletion
duplicate.VersionID, duplicate.DeleteMarkerVersionID = "", version
duplicate.ReplicationState.PurgeTargets = map[string]VersionPurgeStatusType{failing.arn: replication.VersionPurgePending}
duplicate.ReplicationState.VersionPurgeStatusInternal = ""
got := replicateDeleteToTarget(ctx, duplicate, &TargetClient{Client: failing.client, ARN: failing.arn, Bucket: failing.bucket})
if got.VersionPurgeStatus != replication.VersionPurgeComplete {
t.Fatalf("duplicate purge: %+v", got)
}
if oi, err := obj.GetObjectInfo(ctx, failing.bucket, name, ObjectOptions{VersionID: version}); !isErrVersionNotFound(err) && !isErrObjectNotFound(err) {
t.Fatalf("duplicate recreated marker: %+v %v", oi, err)
}
}
if len(p.mrfSaveCh) != 0 || len(p.workers[0]) != 0 {
t.Fatalf("completion left queued work: mrf=%d worker=%d", len(p.mrfSaveCh), len(p.workers[0]))
}
if !tc.creation && !tc.exhaust && !tc.lostReply {
stats.RLock()
got := stats.Cache[bucket].Stats[failing.arn].ReplicatedCount
size := stats.Cache[bucket].Stats[failing.arn].ReplicatedSize
stats.RUnlock()
if got != 1 || size != 0 {
t.Fatalf("purge completion statistic=%d want 1", got)
}
}
t.Logf("%s: %s recovered; %d target(s), persisted MRF, source/target metadata checked", backend, tc.name, tc.targets)
}
func TestReplicationDeleteQueueFullRetryBudget(t *testing.T) {
stats := NewReplicationStats(t.Context(), nil)
p := &ReplicationPool{ctx: t.Context(), objLayer: &replicationMRFTestObjectLayer{}, workers: []chan ReplicationWorkerOperation{make(chan ReplicationWorkerOperation)}, stats: stats, mrfSaveCh: make(chan MRFReplicateEntry, 1), priority: "slow"}
d := DeletedObjectReplicationInfo{Bucket: "bucket", DeletedObject: DeletedObject{ObjectName: "marker", VersionID: mustGetUUID()}, RetryCount: 2}
p.queueReplicaDeleteTask(d)
if got := <-p.mrfSaveCh; got.RetryCount != 3 {
t.Fatalf("queue-full retry count=%d", got.RetryCount)
}
d.RetryCount = mrfRetryLimit
p.queueReplicaDeleteTask(d)
if len(p.mrfSaveCh) != 0 || atomic.LoadUint64(&stats.mrfStats.TotalDroppedCount) != 1 {
t.Fatal("queue-full retry exceeded budget without a visible drop")
}
}
type markerLockFailureLayer struct{ ObjectLayer }
func (markerLockFailureLayer) NewNSLock(string, ...string) RWLocker { return markerFailedLock{} }
type markerFailedLock struct{ RWLocker }
func (markerFailedLock) GetLock(context.Context, *dynamicTimeout) (LockContext, error) {
return LockContext{}, context.DeadlineExceeded
}
type markerLookupLayer struct {
ObjectLayer
mutate func(ObjectInfo, error) (ObjectInfo, error)
lookedUp chan struct{}
}
func (l markerLookupLayer) GetObjectInfo(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
oi, err := l.ObjectLayer.GetObjectInfo(ctx, bucket, object, opts)
defer close(l.lookedUp)
return l.mutate(oi, err)
}
func testReplicationMRFInvalidLookups(ctx context.Context, t *testing.T, obj ObjectLayer, p *ReplicationPool, newPool func() *ReplicationPool, d DeletedObjectReplicationInfo) {
for _, tc := range []struct {
name string
mutate func(ObjectInfo, error) (ObjectInfo, error)
}{
{"empty-info", func(_ ObjectInfo, e error) (ObjectInfo, error) { return ObjectInfo{}, e }},
{"not-a-marker", func(o ObjectInfo, e error) (ObjectInfo, error) { o.DeleteMarker = false; return o, e }},
{"wrong-version", func(o ObjectInfo, e error) (ObjectInfo, error) { o.VersionID = mustGetUUID(); return o, e }},
{"wrong-bucket", func(o ObjectInfo, e error) (ObjectInfo, error) { o.Bucket = "different-bucket"; return o, e }},
{"wrong-object", func(o ObjectInfo, e error) (ObjectInfo, error) { o.Name = "different-object"; return o, e }},
{"zero-modtime", func(o ObjectInfo, e error) (ObjectInfo, error) { o.ModTime = time.Time{}; return o, e }},
{"missing", func(o ObjectInfo, _ error) (ObjectInfo, error) {
return o, ObjectNotFound{Bucket: o.Bucket, Object: o.Name}
}},
{"read-error", func(o ObjectInfo, _ error) (ObjectInfo, error) { return o, InsufficientReadQuorum{} }},
} {
t.Run(tc.name, func(t *testing.T) {
entry := d.ToMRFEntry()
p.saveMRFEntries(ctx, map[string]MRFReplicateEntry{entry.versionID: entry})
fresh := newPool()
looked := make(chan struct{})
fresh.objLayer = markerLookupLayer{ObjectLayer: obj, mutate: tc.mutate, lookedUp: looked}
if err := fresh.queueMRFHeal(); err != nil {
t.Fatal(err)
}
select {
case <-looked:
case <-time.After(3 * time.Second):
t.Fatal("disk MRF entry was not read")
}
select {
case op := <-fresh.workers[0]:
t.Fatalf("invalid lookup scheduled %T", op)
case <-time.After(100 * time.Millisecond):
}
})
}
}
// Capture the actual internal audit event through the supported webhook sink.
func replicationTestAudit(ctx context.Context, t *testing.T, bucket string) (<-chan string, func()) {
t.Helper()
if len(logger.AuditTargets()) != 0 {
t.Fatal("unexpected pre-existing test audit targets")
}
statuses := make(chan string, 16)
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var entry struct {
API struct{ Name, Bucket, Status string }
}
if err := json.NewDecoder(r.Body).Decode(&entry); err != nil {
t.Error(err)
} else if entry.API.Name == ReplicateDeleteAPI && entry.API.Bucket == bucket {
statuses <- entry.API.Status
}
w.WriteHeader(http.StatusOK)
}))
endpoint, err := xnet.ParseHTTPURL(server.URL)
if err != nil {
t.Fatal(err)
}
if errs := logger.UpdateAuditWebhooks(ctx, map[string]loghttp.Config{"r6": {Enabled: true, Name: "r6", Endpoint: endpoint, BatchSize: 1, QueueSize: 128, MaxRetry: 1, RetryIntvl: time.Millisecond, HTTPTimeout: time.Second}}); len(errs) > 0 {
t.Fatal(errs)
}
targets := logger.AuditTargets()
return statuses, func() {
logger.UpdateAuditWebhooks(ctx, nil)
for _, target := range targets {
target.Cancel()
}
server.Close()
}
}
type markerPurgeUpdateLayer struct {
ObjectLayer
t *testing.T
}
func (l markerPurgeUpdateLayer) DeleteObject(ctx context.Context, bucket, object string, opts ObjectOptions) (ObjectInfo, error) {
l.t.Helper()
rs := opts.DeleteReplication
if rs.ReplicationStatusInternal != "" || rs.Targets != nil || rs.ReplicaStatus != "" || !rs.CompositeReplicationStatus().Empty() {
l.t.Fatalf("purge has a nonempty creation update: %+v", rs)
}
if rs.CompositeVersionPurgeStatus().Empty() {
l.t.Fatal("purge update lost its purge status")
}
return l.ObjectLayer.DeleteObject(ctx, bucket, object, opts)
}
+215
View File
@@ -0,0 +1,215 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"fmt"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
)
// Purge uses the version-delete wire form regardless of the in-memory shape.
// Creation status, purge status and successful resync markers are independent.
func TestReplicateDeleteOperationExits(t *testing.T) {
for _, shape := range []string{"marker-creation", "legacy-marker-purge", "marker-purge", "object-purge"} {
purge := shape != "marker-creation"
for _, tc := range []struct {
name string
creation replication.StatusType
priorPurge VersionPurgeStatusType
headCode, deleteCode int
headError string
offline, resync bool
}{
{name: "pending-success", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 204},
{name: "completed-creation", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204},
{name: "existing-marker", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204},
// Use the quorum S3 code without 503, which is classified as backend-down first.
{name: "head-read-quorum", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 400, headError: "SlowDownRead", deleteCode: 204},
{name: "replica-creation-status", creation: replication.Replica, priorPurge: replication.VersionPurgeFailed, headCode: 405, deleteCode: 204},
{name: "head-unavailable", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 503, deleteCode: 204},
{name: "head-forbidden", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 403, deleteCode: 204},
{name: "delete-forbidden", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 403},
{name: "delete-method-rejected", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 405},
{name: "delete-unavailable", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 503},
{name: "offline", creation: replication.Pending, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 204, offline: true},
{name: "retry-failed", creation: replication.Completed, priorPurge: replication.VersionPurgeFailed, headCode: 404, deleteCode: 204},
{name: "purge-complete", creation: replication.Pending, priorPurge: replication.VersionPurgeComplete, headCode: 404, deleteCode: 403},
{name: "resync-success", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 405, deleteCode: 204, resync: true},
{name: "resync-already-purged", creation: replication.Pending, priorPurge: replication.VersionPurgeComplete, headCode: 404, deleteCode: 403, resync: true},
{name: "resync-failure", creation: replication.Completed, priorPurge: replication.VersionPurgePending, headCode: 404, deleteCode: 403, resync: true},
} {
t.Run(shape+"/"+tc.name, func(t *testing.T) {
version := mustGetUUID()
var heads, deletes atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Query().Get("versionId") != version {
t.Errorf("wrong request version: %s", r.URL)
}
switch r.Method {
case http.MethodHead:
heads.Add(1)
w.Header().Set(xhttp.AmzVersionID, version)
w.Header().Set(xhttp.LastModified, time.Now().UTC().Format(http.TimeFormat))
if tc.headCode == 405 {
w.Header().Set(xhttp.AmzDeleteMarker, "true")
}
if tc.headError != "" {
w.Header().Set("x-minio-error-code", tc.headError)
}
w.WriteHeader(tc.headCode)
case http.MethodDelete:
deletes.Add(1)
if got := r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true"; got == purge {
t.Errorf("source delete-marker header=%v, purge=%v", got, purge)
}
w.WriteHeader(tc.deleteCode)
if tc.deleteCode != 204 {
fmt.Fprintf(w, `<Error><Code>%s</Code><Message>injected rejection</Message></Error>`, map[int]string{403: "AccessDenied", 405: "MethodNotAllowed", 503: "ServiceUnavailable"}[tc.deleteCode])
}
default:
t.Errorf("unexpected method %s", r.Method)
}
}))
defer server.Close()
client, err := minio.New(strings.TrimPrefix(server.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
old := globalBucketTargetSys
globalBucketTargetSys = &BucketTargetSys{hc: map[string]epHealth{client.EndpointURL().Host: {Online: !tc.offline}}}
defer func() { globalBucketTargetSys = old }()
d := DeletedObjectReplicationInfo{Bucket: "source", DeletedObject: DeletedObject{ObjectName: "marker", DeleteMarker: shape != "object-purge"}}
if shape == "marker-creation" || shape == "legacy-marker-purge" {
d.DeleteMarkerVersionID = version
} else {
d.VersionID = version
}
d.ReplicationState.Targets = map[string]replication.StatusType{"arn1": tc.creation}
d.ReplicationState.ResetStatusesMap = map[string]string{"arn1": "previous-reset"}
if purge {
d.ReplicationState.PurgeTargets = map[string]VersionPurgeStatusType{"arn1": tc.priorPurge}
}
if tc.resync {
d.OpType = replication.ExistingObjectReplicationType
}
got := replicateDeleteToTarget(t.Context(), d, &TargetClient{Client: client, ARN: "arn1", Bucket: "target", ResetID: "current-reset"})
var wantCreation replication.StatusType
var wantPurge VersionPurgeStatusType
var wantHeads, wantDeletes int32
success := false
if purge {
wantCreation = ""
wantPurge = replication.VersionPurgeComplete
switch {
case tc.priorPurge == replication.VersionPurgeComplete:
success = true
case tc.offline:
wantPurge = replication.VersionPurgeFailed
default:
wantDeletes = 1
success = tc.deleteCode == 204
if !success {
wantPurge = replication.VersionPurgeFailed
}
}
} else {
switch {
case tc.creation == replication.Completed && !tc.resync:
success = true
case tc.offline:
default:
wantHeads = 1
switch tc.headCode {
case 405:
success = true
case 403, 503:
default:
wantDeletes = 1
success = tc.deleteCode == 204
}
}
if success {
wantCreation = replication.Completed
} else {
wantCreation = replication.Failed
}
}
if got.PrevReplicationStatus != tc.creation {
t.Error("previous creation state changed")
}
if got.ReplicationStatus != wantCreation || got.VersionPurgeStatus != wantPurge {
t.Errorf("status=%+v, want creation=%s purge=%s", got, wantCreation, wantPurge)
}
if heads.Load() != wantHeads || deletes.Load() != wantDeletes {
t.Errorf("HEAD/DELETE=%d/%d, want %d/%d", heads.Load(), deletes.Load(), wantHeads, wantDeletes)
}
if (got.Err != nil) == success {
t.Errorf("success=%v, error=%v", success, got.Err)
}
if tc.resync && success {
if !strings.HasSuffix(got.ResyncTimestamp, ";current-reset") {
t.Errorf("successful resync missing reset: %q", got.ResyncTimestamp)
}
} else if got.ResyncTimestamp != "previous-reset" {
t.Errorf("unexpected reset: %q", got.ResyncTimestamp)
}
})
}
}
}
func TestReplicateDeletePurgeMissingTargetState(t *testing.T) {
// Classification belongs to the operation, even if this target has no
// previous purge entry (another target supplies the tracked purge state).
var deletes atomic.Int32
version := mustGetUUID()
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodDelete || r.URL.Query().Get("versionId") != version || r.Header.Get(xhttp.MinIOSourceDeleteMarker) == "true" {
t.Errorf("wrong purge request: %s %s %v", r.Method, r.URL, r.Header)
}
deletes.Add(1)
w.WriteHeader(http.StatusNoContent)
}))
defer server.Close()
client, err := minio.New(strings.TrimPrefix(server.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
old := globalBucketTargetSys
globalBucketTargetSys = &BucketTargetSys{hc: map[string]epHealth{client.EndpointURL().Host: {Online: true}}}
defer func() { globalBucketTargetSys = old }()
d := DeletedObjectReplicationInfo{Bucket: "source", DeletedObject: DeletedObject{ObjectName: "marker", DeleteMarker: true, DeleteMarkerVersionID: version, ReplicationState: ReplicationState{Targets: map[string]replication.StatusType{"arn1": replication.Completed}, PurgeTargets: map[string]VersionPurgeStatusType{"arn2": replication.VersionPurgePending}}}}
got := replicateDeleteToTarget(t.Context(), d, &TargetClient{Client: client, ARN: "arn1", Bucket: "target", ResetID: "reset"})
if deletes.Load() != 1 || got.VersionPurgeStatus != replication.VersionPurgeComplete || got.ReplicationStatus != "" {
t.Fatalf("operation misclassified: %+v, deletes=%d", got, deletes.Load())
}
got.ResyncTimestamp = "resync;reset"
state := getReplicationState(replicatedInfos{Targets: []replicatedTargetInfo{got}}, ReplicationState{}, "")
if state.ResetStatusesMap[targetResetHeader("arn1")] != got.ResyncTimestamp {
t.Fatal("resync timestamp not preserved with nil reset map")
}
}
+121
View File
@@ -0,0 +1,121 @@
// Copyright (c) 2026 PGSTY
// SPDX-License-Identifier: AGPL-3.0-or-later
package cmd
import (
"maps"
"net/http"
"net/textproto"
"reflect"
"strings"
"testing"
xhttp "github.com/minio/minio/internal/http"
)
func TestExtractReplicationMetadataPreservesNormalizedMetadata(t *testing.T) {
for _, tc := range []struct {
name string
wire []string
want string
}{
{name: "absent"},
{name: "transport-only", wire: []string{"aws-chunked"}},
{name: "mixed", wire: []string{"aws-chunked,gzip"}, want: "gzip"},
{name: "gzip", wire: []string{"gzip"}, want: "gzip"},
{name: "transport-last", wire: []string{"gzip,aws-chunked"}, want: "gzip"},
{name: "multiple-values", wire: []string{"aws-chunked", "gzip"}, want: "gzip"},
// Preserve the existing exact-token grammar; whitespace is not normalized here.
{name: "space-before-gzip", wire: []string{"aws-chunked, gzip"}, want: " gzip"},
{name: "space-before-transport", wire: []string{"gzip, aws-chunked"}, want: "gzip, aws-chunked"},
} {
for _, lowercase := range []bool{false, true} {
name := tc.name + "/canonical"
if lowercase {
name = tc.name + "/lowercase"
}
t.Run(name, func(t *testing.T) {
header := http.Header{
"Content-Type": []string{"application/octet-stream"},
"X-Amz-Meta-Source": []string{"raw"},
"X-Minio-Replication-Server-Side-Encryption-Sealed-Key": []string{"sealed-key"},
"X-Minio-Replication-Server-Side-Encryption-Seal-Algorithm": []string{"DAREv2-HMAC-SHA256"},
"X-Minio-Replication-Server-Side-Encryption-Iv": []string{"iv"},
"X-Minio-Replication-Encrypted-Multipart": []string{""},
"X-Minio-Replication-Actual-Object-Size": []string{"1"},
ReplicationSsecChecksumHeader: []string{"checksum"},
xhttp.AmzMetaUnencryptedContentLength: []string{"injected-length"},
xhttp.AmzMetaUnencryptedContentMD5: []string{"injected-md5"},
}
if tc.wire != nil {
header[xhttp.ContentEncoding] = tc.wire
}
if lowercase {
h := make(http.Header, len(header))
for k, v := range header {
h[strings.ToLower(k)] = v
}
header = h
}
metadata, err := extractMetadata(t.Context(), textproto.MIMEHeader(header))
if err != nil {
t.Fatal(err)
}
if metadata["content-encoding"] != tc.want {
t.Fatalf("ordinary encoding=%q want=%q", metadata["content-encoding"], tc.want)
}
for _, internal := range replicationToInternalHeaders {
if _, ok := metadata[internal]; ok {
t.Fatalf("ordinary request accepted internal field %s", internal)
}
}
// Callers own ordinary metadata and may transform it after extraction.
metadata["content-type"] = "application/wasm"
for k := range metadata {
if strings.EqualFold(k, "x-amz-meta-source") {
metadata[k] = "caller"
}
}
want := maps.Clone(metadata)
maps.Copy(want, map[string]string{
"X-Minio-Internal-Server-Side-Encryption-Sealed-Key": "sealed-key",
"X-Minio-Internal-Server-Side-Encryption-Seal-Algorithm": "DAREv2-HMAC-SHA256",
"X-Minio-Internal-Server-Side-Encryption-Iv": "iv",
"X-Minio-Internal-Encrypted-Multipart": "",
"X-Minio-Internal-Actual-Object-Size": "1",
ReplicationSsecChecksumHeader: "checksum",
})
for range 2 {
if err := extractReplicationMetadataFromMime(t.Context(), textproto.MIMEHeader(header), metadata); err != nil {
t.Fatal(err)
}
if !reflect.DeepEqual(metadata, want) {
t.Errorf("restoration changed normalized metadata: got %#v want %#v", metadata, want)
}
}
if tc.want == "" {
if _, present := metadata["content-encoding"]; present {
t.Error("transport-only content-encoding key restored")
}
}
for _, key := range []string{xhttp.AmzMetaUnencryptedContentLength, xhttp.AmzMetaUnencryptedContentMD5} {
if _, present := caseInsensitiveMap(metadata).Lookup(key); present {
t.Errorf("redacted metadata restored: %s", key)
}
}
})
}
}
}
func TestExtractReplicationMetadataNilHeader(t *testing.T) {
metadata := map[string]string{"content-type": "application/wasm"}
want := maps.Clone(metadata)
if err := extractReplicationMetadataFromMime(t.Context(), nil, metadata); err != errInvalidArgument {
t.Fatalf("nil header: got %v want %v", err, errInvalidArgument)
}
if !reflect.DeepEqual(metadata, want) {
t.Fatalf("nil input changed metadata: %#v", metadata)
}
}
+599
View File
@@ -0,0 +1,599 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"context"
"encoding/xml"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/kms"
)
const r5TagStamp = ReservedMetadataPrefixLower + TaggingTimestamp
func r5Capacity(z *erasureServerPools) func() {
var restores []func()
for _, pool := range z.serverPools {
for _, set := range pool.sets {
old := set.getDisks
disks := append([]StorageAPI(nil), old()...)
for i := range disks {
disks[i] = tagTestCapacityDisk{StorageAPI: disks[i]}
}
set.getDisks = func() []StorageAPI { return disks }
restores = append(restores, func() { set.getDisks = old })
}
}
return func() {
for _, restore := range restores {
restore()
}
}
}
func r5Request(t *testing.T, router http.Handler, cred auth.Credentials, method, path, body string, headers map[string]string) *httptest.ResponseRecorder {
t.Helper()
r, err := newTestSignedRequestV4(method, path, int64(len(body)), strings.NewReader(body), cred.AccessKey, cred.SecretKey, headers)
if err != nil {
t.Fatal(err)
}
w := httptest.NewRecorder()
router.ServeHTTP(w, r)
return w
}
func r5Stored(t *testing.T, obj interface {
GetObjectInfo(context.Context, string, string, ObjectOptions) (ObjectInfo, error)
}, bucket, name, vid, wantTags, wantStamp string,
) ObjectInfo {
t.Helper()
oi, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if err != nil {
t.Fatal(err)
}
if oi.UserTags != wantTags || oi.UserDefined[r5TagStamp] != wantStamp {
t.Fatalf("%s(%s): tags=%q stamp=%q, want %q %q", name, vid, oi.UserTags, oi.UserDefined[r5TagStamp], wantTags, wantStamp)
}
return oi
}
// The request bytes are signed and enter the real API and storage implementation.
// Equal/stale retransmits may retain the existing 412 duplicate response.
func r5Receive(t *testing.T, obj ObjectLayer, router http.Handler, cred auth.Credentials, bucket, operation string, source ObjectInfo, stamp string, afterInit func()) {
t.Helper()
opts, _, err := putReplicationOpts(t.Context(), "", source)
if err != nil {
t.Fatal(err)
}
if operation == "multipart" {
opts.Internal.SourceMTime = time.Time{}
}
headers := make(map[string]string)
for k, vs := range opts.Header() {
if len(vs) > 0 {
headers[k] = vs[0]
}
}
if stamp == "" {
delete(headers, http.CanonicalHeaderKey(xhttp.MinIOSourceTaggingTimestamp))
delete(headers, xhttp.MinIOSourceTaggingTimestamp)
} else {
headers[http.CanonicalHeaderKey(xhttp.MinIOSourceTaggingTimestamp)] = stamp
}
path := "/" + bucket + "/" + source.Name + "?versionId=" + source.VersionID
var w *httptest.ResponseRecorder
switch operation {
case "copy", "copy-default":
maps.Copy(headers, getCopyObjMetadata(source, ""))
headers[xhttp.AmzCopySource] = "/" + bucket + "/" + source.Name + "?versionId=" + source.VersionID
if operation == "copy" {
headers[xhttp.AmzMetadataDirective] = "REPLACE"
}
w = r5Request(t, router, cred, http.MethodPut, path, "", headers)
case "put":
w = r5Request(t, router, cred, http.MethodPut, path, "data", headers)
case "multipart":
w = r5Request(t, router, cred, http.MethodPost, path+"&uploads", "", headers)
if w.Code == http.StatusPreconditionFailed {
return
}
if w.Code != http.StatusOK {
t.Fatalf("init: %d %s", w.Code, w.Body.String())
}
var init struct {
UploadID string `xml:"UploadId"`
}
if err := xml.Unmarshal(w.Body.Bytes(), &init); err != nil || init.UploadID == "" {
t.Fatalf("init XML: %v %s", err, w.Body.String())
}
mi, err := obj.GetMultipartInfo(t.Context(), bucket, source.Name, init.UploadID, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
if stamp != "" && mi.UserDefined[r5TagStamp] != stamp {
t.Fatalf("upload persisted stamp=%q, want %q", mi.UserDefined[r5TagStamp], stamp)
}
partPath := "/" + bucket + "/" + source.Name + "?uploadId=" + url.QueryEscape(init.UploadID)
ph := map[string]string{xhttp.MinIOSourceReplicationRequest: "true"}
part := r5Request(t, router, cred, http.MethodPut, partPath+"&partNumber=1", "data", ph)
if part.Code != http.StatusOK {
t.Fatalf("part: %d %s", part.Code, part.Body.String())
}
if afterInit != nil {
afterInit()
}
body := "<CompleteMultipartUpload><Part><PartNumber>1</PartNumber><ETag>" + canonicalizeETag(part.Header()[xhttp.ETag][0]) + "</ETag></Part></CompleteMultipartUpload>"
ph[xhttp.MinIOSourceMTime] = source.ModTime.Format(time.RFC3339Nano)
ph[xhttp.MinIOSourceETag] = source.ETag
w = r5Request(t, router, cred, http.MethodPost, partPath, body, ph)
}
if w.Code != http.StatusOK && w.Code != http.StatusPreconditionFailed {
t.Fatalf("%s: %d %s", operation, w.Code, w.Body.String())
}
}
func TestAPITaggingReplicationOrdering(t *testing.T) { r5APIOrdering(t, false) }
func TestAPITaggingReplicationOrderingKMS(t *testing.T) { r5APIOrdering(t, true) }
func r5APIOrdering(t *testing.T, encrypted bool) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
if encrypted {
prev := GlobalKMS
GlobalKMS = kms.NewStub("r5-tag-order")
defer func() { GlobalKMS = prev }()
sse := []byte(`<ServerSideEncryptionConfiguration xmlns="http://s3.amazonaws.com/doc/2006-03-01/"><Rule><ApplyServerSideEncryptionByDefault><SSEAlgorithm>aws:kms</SSEAlgorithm><KMSMasterKeyID>r5-tag-order</KMSMasterKeyID></ApplyServerSideEncryptionByDefault></Rule></ServerSideEncryptionConfiguration>`)
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketSSEConfig, sse); err != nil {
t.Fatal(err)
}
}
base := time.Now().UTC().Add(-5 * time.Hour)
for _, op := range []string{"copy", "copy-default", "put", "multipart"} {
for _, version := range []string{"uuid", "null"} {
t.Run(instance+"/"+op+"/"+version, func(t *testing.T) {
name := op + "-" + version
vid := mustGetUUID()
if version == "null" {
vid = nullVersionID
}
original, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true, VersionID: vid, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=original", r5TagStamp: base.Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
if op == "multipart" {
// The production sender uses multipart only for multipart
// sources. Preserve a real multipart ETag and part layout.
metadata := maps.Clone(original.UserDefined)
metadata[xhttp.AmzObjectTagging] = original.UserTags
mp, err := obj.NewMultipartUpload(t.Context(), bucket, name, ObjectOptions{Versioned: true, VersionID: vid, UserDefined: metadata})
if err != nil {
t.Fatal(err)
}
part, err := obj.PutObjectPart(t.Context(), bucket, name, mp.UploadID, 1, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{})
if err != nil {
t.Fatal(err)
}
original, err = obj.CompleteMultipartUpload(t.Context(), bucket, name, mp.UploadID, []CompletePart{{PartNumber: 1, ETag: part.ETag}}, ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
}
// A later unrelated version must not contribute tags to an explicitly addressed UUID/null version.
latest, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("other"), 5, "", ""), ObjectOptions{Versioned: true, UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=latest", r5TagStamp: base.Add(10 * time.Hour).Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
for _, event := range []struct {
name, tags string
hours int
wantTags string
wantHours int
}{
{"delete", "", 3, "", 3}, {"stale", "key=stale", 2, "", 3}, {"equal-conflict", "key=conflict", 3, "", 3}, {"newer", "key=new", 4, "key=new", 4},
} {
t.Run(event.name, func(t *testing.T) {
source := original
source.VersionID = vid
source.UserDefined = maps.Clone(original.UserDefined)
source.UserTags = event.tags
stamp := base.Add(time.Duration(event.hours) * time.Hour).Format(time.RFC3339Nano)
source.UserDefined[r5TagStamp] = stamp
r5Receive(t, obj, router, cred, bucket, op, source, stamp, nil)
r5Stored(t, obj, bucket, name, vid, event.wantTags, base.Add(time.Duration(event.wantHours)*time.Hour).Format(time.RFC3339Nano))
r5Stored(t, obj, bucket, name, latest.VersionID, "key=latest", base.Add(10*time.Hour).Format(time.RFC3339Nano))
})
}
source := original
source.VersionID = vid
source.UserTags = "key=unversioned-event"
r5Receive(t, obj, router, cred, bucket, op, source, "", nil)
r5Stored(t, obj, bucket, name, vid, "key=new", base.Add(4*time.Hour).Format(time.RFC3339Nano))
get := r5Request(t, router, cred, http.MethodGet, "/"+bucket+"/"+name+"?versionId="+vid, "", nil)
if get.Code != http.StatusOK || get.Body.String() != "data" {
t.Fatalf("plaintext GET: %d %q", get.Code, get.Body.String())
}
})
}
}
}})
}
func TestAPITaggingMultipartCommitRechecksRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
base := time.Now().UTC().Add(-time.Hour)
source := ObjectInfo{Name: "commit-recheck", VersionID: mustGetUUID(), ModTime: base, UserTags: "key=incoming", UserDefined: map[string]string{r5TagStamp: base.Format(time.RFC3339Nano)}}
later := base.Add(time.Minute).Format(time.RFC3339Nano)
r5Receive(t, obj, router, cred, bucket, "multipart", source, base.Format(time.RFC3339Nano), func() {
_, err := obj.PutObject(t.Context(), bucket, source.Name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true, VersionID: source.VersionID, MTime: base, UserDefined: map[string]string{r5TagStamp: later}})
if err != nil {
t.Fatal(err)
}
})
r5Stored(t, obj, bucket, source.Name, source.VersionID, "", later)
t.Logf("%s: deletion committed between initiation and completion survived", instance)
}})
}
func TestAPILocalTaggingAlwaysAdvancesRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
oi, err := obj.PutObject(t.Context(), bucket, "local-tags", mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
path := "/" + bucket + "/local-tags?tagging&versionId=" + oi.VersionID
for n, method := range []string{http.MethodPut, http.MethodDelete, http.MethodDelete, http.MethodPut} {
body := ""
want := ""
status := http.StatusNoContent
if method == http.MethodPut {
status = http.StatusOK
body = "<Tagging><TagSet/></Tagging>"
if n == 0 {
want = "key=local"
body = "<Tagging><TagSet><Tag><Key>key</Key><Value>local</Value></Tag></TagSet></Tagging>"
}
}
before := time.Now().UTC()
w := r5Request(t, router, cred, method, path, body, nil)
if w.Code != status {
t.Fatalf("%s: %d %s", method, w.Code, w.Body.String())
}
now, err := obj.GetObjectInfo(t.Context(), bucket, "local-tags", ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
stamp, err := time.Parse(time.RFC3339Nano, now.UserDefined[r5TagStamp])
if err != nil || stamp.Before(before) || now.UserTags != want {
t.Fatalf("local %s: tags=%q timestamp=%q err=%v", method, now.UserTags, now.UserDefined[r5TagStamp], err)
}
if !now.ModTime.Equal(oi.ModTime) {
t.Fatal("tagging changed the object's data modification time")
}
t.Logf("%s mutation %d persisted tags=%q timestamp=%s", instance, n, now.UserTags, stamp)
}
// Ordinary COPY with empty REPLACE has the same local-deletion semantics.
before := time.Now().UTC()
headers := map[string]string{xhttp.AmzCopySource: "/" + bucket + "/local-tags?versionId=" + oi.VersionID, xhttp.AmzMetadataDirective: "REPLACE", xhttp.AmzTagDirective: "REPLACE"}
w := r5Request(t, router, cred, http.MethodPut, "/"+bucket+"/copied-empty", "", headers)
if w.Code != http.StatusOK {
t.Fatalf("copy: %d %s", w.Code, w.Body.String())
}
copied, err := obj.GetObjectInfo(t.Context(), bucket, "copied-empty", ObjectOptions{})
if err != nil {
t.Fatal(err)
}
stamp, err := time.Parse(time.RFC3339Nano, copied.UserDefined[r5TagStamp])
if err != nil || stamp.Before(before) || copied.UserTags != "" {
t.Fatalf("local COPY: %v %+v", err, copied)
}
}})
}
func TestTaggingTimestampWire(t *testing.T) {
now := time.Now().UTC()
for _, tag := range []string{"", "key=value"} {
for _, stamp := range []string{"", now.Format(time.RFC3339Nano), "invalid"} {
t.Run(fmt.Sprintf("%s/%s", tag, stamp), func(t *testing.T) {
source := ObjectInfo{ModTime: now.Add(-time.Hour), UserTags: tag, UserDefined: map[string]string{}}
if stamp != "" {
source.UserDefined[r5TagStamp] = stamp
}
opts, _, err := putReplicationOpts(t.Context(), "", source)
if stamp == "invalid" {
if err == nil {
t.Fatal("invalid timestamp accepted")
}
return
}
if err != nil {
t.Fatal(err)
}
want := stamp
if want == "" && tag != "" {
want = source.ModTime.Format(time.RFC3339Nano)
}
if got := opts.Header().Get(xhttp.MinIOSourceTaggingTimestamp); got != want {
t.Fatalf("wire timestamp=%q want %q", got, want)
}
})
}
}
}
func TestAPIPoolsTaggingReplicaDeletion(t *testing.T) {
z, bucket := consistencyPools(t)
defer r5Capacity(z)()
if err := newTestConfig(globalMinioDefaultRegion, z); err != nil {
t.Fatal(err)
}
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
router := initTestAPIEndPoints(z, nil)
base := time.Now().UTC().Add(-5 * time.Hour)
for _, op := range []string{"copy", "copy-default", "put", "multipart"} {
for _, kind := range []string{"uuid", "null"} {
t.Run(op+"/"+kind, func(t *testing.T) {
vid := mustGetUUID()
if kind == "null" {
vid = nullVersionID
}
name := "pool-tags-" + op + "-" + kind
var source ObjectInfo
for pool := range 2 {
tag := "key=stale-pool"
ts := base
if pool == 1 {
tag = ""
ts = base.Add(3 * time.Hour)
}
source = putConsistencyObject(t, z, bucket, name, pool, "data", ObjectOptions{Versioned: true, VersionID: vid, MTime: base, UserDefined: map[string]string{xhttp.AmzObjectTagging: tag, r5TagStamp: ts.Format(time.RFC3339Nano)}})
}
// Clear all copies through the signed local handler, then replay a stale incoming state.
req := r5Request(t, router, globalActiveCred, http.MethodDelete, "/"+bucket+"/"+name+"?tagging&versionId="+vid, "", nil)
if req.Code != http.StatusNoContent {
t.Fatalf("DELETE: %d %s", req.Code, req.Body.String())
}
current, err := z.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if err != nil {
t.Fatal(err)
}
deletedAt := current.UserDefined[r5TagStamp]
for pool := range 2 {
r5Stored(t, z.serverPools[pool], bucket, name, vid, "", deletedAt)
}
source.VersionID = vid
source.UserTags = "key=delayed"
source.UserDefined = map[string]string{r5TagStamp: base.Add(time.Hour).Format(time.RFC3339Nano)}
r5Receive(t, z, router, globalActiveCred, bucket, op, source, source.UserDefined[r5TagStamp], nil)
// The addressed version must remain readable through normal routing.
r5Stored(t, z, bucket, name, vid, "", deletedAt)
// Existing duplicate suppression may leave both identical tombstones; any retained copy must be correct.
for pool := range 2 {
got, err := z.serverPools[pool].GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: vid})
if isErrVersionNotFound(err) {
continue
}
if err != nil {
t.Fatal(err)
}
if got.UserTags != "" || got.UserDefined[r5TagStamp] != deletedAt {
t.Fatalf("pool %d: tags=%q stamp=%q want deletion %q", pool, got.UserTags, got.UserDefined[r5TagStamp], deletedAt)
}
}
})
}
}
}
func TestAPITaggingSSECRotationPreservesDeletionRevision(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
oldTLS := globalIsTLS
globalIsTLS = true
defer func() { globalIsTLS = oldTLS }()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
keys := [][]byte{[]byte(strings.Repeat("a", 32)), []byte(strings.Repeat("b", 32)), []byte(strings.Repeat("c", 32)), []byte(strings.Repeat("d", 32))}
const name = "tag-rotation"
putCopyChecksumSource(t, router, cred, bucket, name, []byte("data"), ssecKeyHeaders(keys[0], false))
oi, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{})
if err != nil {
t.Fatal(err)
}
base := time.Now().UTC().Add(-time.Hour)
for i, event := range []struct {
tags string
delta int
wantTags string
wantDelta int
}{
{"key=live", 1, "key=live", 1}, {"", 3, "", 3}, {"key=stale", 2, "", 3},
} {
h := ssecKeyHeaders(keys[i], true)
maps.Copy(h, ssecKeyHeaders(keys[i+1], false))
h[xhttp.AmzObjectTagging] = event.tags
h[xhttp.AmzTagDirective] = "REPLACE"
h[xhttp.MinIOSourceTaggingTimestamp] = base.Add(time.Duration(event.delta) * time.Minute).Format(time.RFC3339Nano)
sendReplicaLockCopy(t, router, cred, bucket, name, oi.VersionID, h)
r5Stored(t, obj, bucket, name, oi.VersionID, event.wantTags, base.Add(time.Duration(event.wantDelta)*time.Minute).Format(time.RFC3339Nano))
}
get := r5Request(t, router, cred, http.MethodGet, "/"+bucket+"/"+name+"?versionId="+oi.VersionID, "", ssecKeyHeaders(keys[3], false))
if get.Code != http.StatusOK || get.Body.String() != "data" {
t.Fatalf("%s GET after rotations: %d %q", instance, get.Code, get.Body.String())
}
}})
}
func TestTaggingRepeatedValueNeedsRevisionDelivery(t *testing.T) {
for _, tag := range []string{"", "key=same"} {
now := time.Now().UTC()
source := ObjectInfo{ModTime: now, UserTags: tag, UserDefined: map[string]string{r5TagStamp: now.Add(time.Hour).Format(time.RFC3339Nano)}}
target := minio.ObjectInfo{LastModified: now}
if tag != "" {
target.UserTags = map[string]string{"key": "same"}
target.UserTagCount = 1
}
if got := getReplicationAction(source, target, replication.MetadataReplicationType); got != replicateMetadata {
t.Errorf("equal tags %q suppress a newer revision: got %s; an intervening delayed deletion can win", tag, got)
}
}
}
func TestTaggingProductionCopyWireShape(t *testing.T) {
got := make(chan http.Header, 1)
peer := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
got <- r.Header.Clone()
w.Header().Set(xhttp.ContentType, "application/xml")
w.Write([]byte(`<CopyObjectResult><LastModified>2026-09-15T01:00:00Z</LastModified><ETag>"abc"</ETag></CopyObjectResult>`))
}))
defer peer.Close()
c, err := minio.New(strings.TrimPrefix(peer.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
now := time.Now().UTC()
oi := ObjectInfo{Name: "object", ModTime: now, ETag: "abc", VersionID: mustGetUUID()}
core := minio.Core{Client: c}
_, err = core.CopyObject(t.Context(), "bucket", "object", "bucket", "object", getCopyObjMetadata(oi, ""), minio.CopySrcOptions{VersionID: oi.VersionID}, minio.PutObjectOptions{Internal: minio.AdvancedPutOptions{SourceVersionID: oi.VersionID, ReplicationRequest: true, TaggingTimestamp: now}})
if err != nil {
t.Fatal(err)
}
h := <-got
t.Logf("SDK wire: metadata-directive=%q tagging-directive=%q tagging=%q time=%q", h.Get(xhttp.AmzMetadataDirective), h.Get(xhttp.AmzTagDirective), h.Get(xhttp.AmzObjectTagging), h.Get(xhttp.MinIOSourceTaggingTimestamp))
if h.Get(xhttp.AmzMetadataDirective) != "" || h.Get(xhttp.AmzTagDirective) != "REPLACE" || h.Get(xhttp.MinIOSourceTaggingTimestamp) != now.Format(time.RFC3339Nano) {
t.Fatalf("unexpected SDK wire: %v", h)
}
}
func TestLocalTaggingCommitCannotRegressRevision(t *testing.T) {
z, bucket := consistencyPools(t)
incoming := "2026-09-15T01:00:00Z"
newer := "2026-09-15T02:00:00Z"
newest := "2026-09-15T03:00:00Z"
vid := mustGetUUID()
name := "inverted-local-tags"
for pool := range 2 {
putConsistencyObject(t, z, bucket, name, pool, "data", ObjectOptions{Versioned: true, VersionID: vid, UserDefined: map[string]string{r5TagStamp: []string{newer, newest}[pool]}})
}
oi, err := z.PutObjectTags(t.Context(), bucket, name, "key=after-delete", ObjectOptions{VersionID: vid, UserDefined: map[string]string{r5TagStamp: incoming}})
if err != nil {
t.Fatal(err)
}
want := "2026-09-15T03:00:00.000000001Z"
for pool := range 2 {
r5Stored(t, z.serverPools[pool], bucket, name, vid, "key=after-delete", want)
}
if oi.UserDefined[r5TagStamp] != want {
t.Fatalf("response stamp=%q want %q", oi.UserDefined[r5TagStamp], want)
}
// The single-set guard also applies when the pools dispatcher is bypassed.
_, err = z.serverPools[0].PutObjectTags(t.Context(), bucket, name, "", ObjectOptions{VersionID: vid, UserDefined: map[string]string{r5TagStamp: incoming}})
if err != nil {
t.Fatal(err)
}
r5Stored(t, z.serverPools[0], bucket, name, vid, "", "2026-09-15T03:00:00.000000002Z")
}
func TestTaggingReplicaContentDuplicateGuard(t *testing.T) {
stamp := time.Date(2026, 9, 15, 1, 0, 0, 0, time.UTC)
for _, tc := range []struct {
name string
delta int
stored string
trusted, replica bool
ifMatch, ifNone string
wantSkip bool
}{
{name: "newer", delta: 1, trusted: true, replica: true},
{name: "equal", trusted: true, replica: true, wantSkip: true},
{name: "older", delta: -1, trusted: true, replica: true, wantSkip: true},
{name: "invalid-stored", stored: "invalid", delta: 1, trusted: true, replica: true},
{name: "untrusted", delta: 1, wantSkip: true},
{name: "marker-only", delta: 1, trusted: true, wantSkip: true},
{name: "if-match-fails", delta: 1, trusted: true, replica: true, ifMatch: "other", wantSkip: true},
{name: "if-none-match-fails", delta: 1, trusted: true, replica: true, ifNone: "etag", wantSkip: true},
{name: "if-match-passes", delta: 1, trusted: true, replica: true, ifMatch: "etag"},
} {
t.Run(tc.name, func(t *testing.T) {
ctx := withReplicationTrust(t.Context(), tc.trusted, tc.replica)
r := httptest.NewRequest(http.MethodPut, "/bucket/object", nil).WithContext(ctx)
if tc.ifMatch != "" {
r.Header.Set(xhttp.IfMatch, tc.ifMatch)
}
if tc.ifNone != "" {
r.Header.Set(xhttp.IfNoneMatch, tc.ifNone)
}
stored := tc.stored
if stored == "" {
stored = stamp.Format(time.RFC3339Nano)
}
oi := ObjectInfo{ModTime: stamp, VersionID: mustGetUUID(), ETag: "etag", UserDefined: map[string]string{r5TagStamp: stored}}
opts := ObjectOptions{VersionID: oi.VersionID, PreserveETag: oi.ETag, ReplicationRequest: tc.trusted, ReplicationSourceTaggingTimestamp: stamp.Add(time.Duration(tc.delta) * time.Second)}
if skip := checkPreconditionsPUT(ctx, httptest.NewRecorder(), r, oi, opts); skip != tc.wantSkip {
t.Fatalf("skip=%v, want %v", skip, tc.wantSkip)
}
})
}
}
func TestAPITaggingUnqualifiedCopyOrdering(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
stamp := time.Now().UTC().Add(-time.Hour)
oi, err := obj.PutObject(t.Context(), bucket, "unqualified-tags", mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{UserDefined: map[string]string{xhttp.AmzObjectTagging: "key=stored", r5TagStamp: stamp.Format(time.RFC3339Nano)}})
if err != nil {
t.Fatal(err)
}
source := oi
source.UserTags = ""
source.UserDefined = maps.Clone(oi.UserDefined)
r5Receive(t, obj, router, cred, bucket, "copy", source, stamp.Format(time.RFC3339Nano), nil)
r5Stored(t, obj, bucket, oi.Name, "", "key=stored", stamp.Format(time.RFC3339Nano))
later := stamp.Add(time.Minute).Format(time.RFC3339Nano)
r5Receive(t, obj, router, cred, bucket, "copy-default", source, later, nil)
r5Stored(t, obj, bucket, oi.Name, "", "", later)
source.UserTags = "key=delayed"
r5Receive(t, obj, router, cred, bucket, "copy-default", source, stamp.Format(time.RFC3339Nano), nil)
r5Stored(t, obj, bucket, oi.Name, "", "", later)
t.Logf("%s: unqualified COPY keeps stored ties and ordered deletion", instance)
}})
}
+154
View File
@@ -0,0 +1,154 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/minio/madmin-go/v3"
"github.com/minio/minio-go/v7"
"github.com/minio/minio/internal/auth"
"github.com/minio/minio/internal/bucket/replication"
xhttp "github.com/minio/minio/internal/http"
"github.com/minio/minio/internal/once"
)
func r5ReplicationFixture(t *testing.T, obj ObjectLayer, bucket string, client *minio.Client) (chan ReplicationWorkerOperation, func()) {
t.Helper()
const arn = "arn:minio:replication::af470089-d354-4473-934c-9e1f52f6da89:bucket"
target := &TargetClient{Client: client, ARN: arn, Bucket: bucket}
globalBucketTargetSys.arnRemotesMap[arn] = arnTarget{Client: target, lastRefresh: UTCNow()}
globalBucketTargetSys.targetsMap[bucket] = []madmin.BucketTarget{{Arn: arn, TargetBucket: bucket}}
meta, err := globalBucketMetadataSys.Get(bucket)
if err != nil {
t.Fatal(err)
}
cfg := configs[0]
cfg.RoleArn = arn
meta.replicationConfig = &cfg
globalBucketMetadataSys.Set(bucket, meta)
worker := make(chan ReplicationWorkerOperation, 10)
previous := globalReplicationPool
globalReplicationPool = once.NewSingleton[ReplicationPool]()
globalReplicationPool.Set(&ReplicationPool{ctx: t.Context(), objLayer: obj, workers: []chan ReplicationWorkerOperation{worker}, stats: globalReplicationStats.Load(), mrfSaveCh: make(chan MRFReplicateEntry, 10)})
return worker, func() { globalReplicationPool = previous }
}
func TestTaggingReplicationSenderRetryAndAcknowledgment(t *testing.T) {
ExecObjectLayerAPITest(ExecObjectLayerAPITestArgs{t: t, objAPITest: func(obj ObjectLayer, instance, bucket string, router http.Handler, cred auth.Credentials, t *testing.T) {
defer r5Capacity(obj.(*erasureServerPools))()
if _, err := globalBucketMetadataSys.Update(t.Context(), bucket, bucketVersioningConfig, enabledBucketVersioningConfig); err != nil {
t.Fatal(err)
}
const name = "tagging-old-queue"
oi, err := obj.PutObject(t.Context(), bucket, name, mustGetPutObjReader(t, strings.NewReader("data"), 4, "", ""), ObjectOptions{Versioned: true})
if err != nil {
t.Fatal(err)
}
requests := make(chan http.Header, 4)
var attempts atomic.Int32
peer := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set(xhttp.AmzVersionID, oi.VersionID)
w.Header().Set(xhttp.ETag, "\""+oi.ETag+"\"")
w.Header().Set(xhttp.LastModified, oi.ModTime.Format(http.TimeFormat))
w.Header().Set(xhttp.ContentType, oi.ContentType)
if r.Method == http.MethodHead {
w.Header().Set(xhttp.ContentLength, "4")
w.WriteHeader(http.StatusOK)
return
}
requests <- r.Header.Clone()
w.Header().Set(xhttp.ContentType, "application/xml")
if attempts.Add(1) == 1 {
w.WriteHeader(http.StatusServiceUnavailable)
w.Write([]byte(`<Error><Code>SlowDown</Code><Message>retry fixture</Message></Error>`))
return
}
w.Write([]byte("<CopyObjectResult><LastModified>" + oi.ModTime.Format(time.RFC3339Nano) + "</LastModified><ETag>\"" + oi.ETag + "\"</ETag></CopyObjectResult>"))
}))
defer peer.Close()
client, err := minio.New(strings.TrimPrefix(peer.URL, "http://"), &minio.Options{Region: "us-east-1", MaxRetries: 1})
if err != nil {
t.Fatal(err)
}
worker, cleanup := r5ReplicationFixture(t, obj, bucket, client)
defer cleanup()
body := `<Tagging><TagSet><Tag><Key>key</Key><Value>queued</Value></Tag></TagSet></Tagging>`
w := r5Request(t, router, cred, http.MethodPut, "/"+bucket+"/"+name+"?tagging&versionId="+oi.VersionID, body, nil)
if w.Code != http.StatusOK || len(worker) != 1 {
t.Fatalf("tagging PUT: %d %s queued=%d", w.Code, w.Body.String(), len(worker))
}
old := (<-worker).(ReplicateObjectInfo)
// Delete through the actual handler before processing the old task.
w = r5Request(t, router, cred, http.MethodDelete, "/"+bucket+"/"+name+"?tagging&versionId="+oi.VersionID, "", nil)
if w.Code != http.StatusNoContent || len(worker) != 1 {
t.Fatalf("tagging DELETE: %d %s queued=%d", w.Code, w.Body.String(), len(worker))
}
deleted, err := obj.GetObjectInfo(t.Context(), bucket, name, ObjectOptions{VersionID: oi.VersionID})
if err != nil {
t.Fatal(err)
}
stamp := deleted.UserDefined[r5TagStamp]
for attempt := 0; attempt < 2; attempt++ {
result := replicateObject(t.Context(), old, obj)
want := replication.Failed
if attempt == 1 {
want = replication.Completed
}
if result.ReplicationStatus() != want {
t.Fatalf("attempt %d result=%+v want %s", attempt, result, want)
}
if len(result.Targets) != 1 || result.Targets[0].ReplicationAction != replicateMetadata || (result.Targets[0].Err != nil) != (attempt == 0) {
t.Fatalf("attempt %d reported wrong action/error: %+v", attempt, result)
}
r5Stored(t, obj, bucket, name, oi.VersionID, "", stamp)
select {
case h := <-requests:
if h.Get(xhttp.MinIOSourceTaggingTimestamp) != stamp || h.Get(xhttp.AmzObjectTagging) != "" || h.Get(xhttp.AmzTagDirective) != "REPLACE" || h.Get(xhttp.AmzMetadataDirective) != "" {
t.Fatalf("sender did not carry current deletion: %v", h)
}
t.Logf("%s attempt %d sent tags=%q timestamp=%s status=%s", instance, attempt, h.Get(xhttp.AmzObjectTagging), stamp, want)
default:
t.Fatal("no metadata COPY sent for same-empty target")
}
}
// With a real outgoing rule enabled, a signed incoming replica COPY
// must not queue another outgoing event and create a feedback loop.
before := len(worker)
r5Receive(t, obj, router, cred, bucket, "copy", deleted, stamp, nil)
if len(worker) != before {
t.Fatal("incoming replica COPY scheduled another outgoing event")
}
// The metadata sender must fail malformed stored revisions before COPY,
// just as the full retransmission option builder does.
_, err = obj.PutObjectTags(t.Context(), bucket, name, "", ObjectOptions{VersionID: oi.VersionID, UserDefined: map[string]string{r5TagStamp: "invalid"}})
if err != nil {
t.Fatal(err)
}
target := globalBucketTargetSys.GetRemoteTargetClient(bucket, globalBucketTargetSys.targetsMap[bucket][0].Arn)
invalid := old.replicateAll(t.Context(), obj, target)
if invalid.ReplicationStatus != replication.Failed || invalid.Err == nil || len(requests) != 0 {
t.Fatalf("invalid timestamp was not rejected before send: %+v", invalid)
}
}})
}
+2
View File
@@ -901,6 +901,8 @@ func serverMain(ctx *cli.Context) {
close(globalGridStart)
close(globalLockGridStart)
// The HTTP/1 listener preserves absolute header deadlines and renews the
// body read/write idle limits, so transfers may outlast IdleTimeout.
httpServer := xhttp.NewServer(getServerListenAddrs()).
UseHandler(setCriticalErrorHandler(corsHandler(handler))).
UseTLSConfig(newTLSConfig(getCert)).
+97
View File
@@ -0,0 +1,97 @@
// Copyright (c) 2026 Feng Ruohang
//
// This file is part of Silo Object Storage stack
//
// This program is free software: you can redistribute it and/or modify
// it under the terms of the GNU Affero General Public License as published by
// the Free Software Foundation, either version 3 of the License, or
// (at your option) any later version.
//
// This program is distributed in the hope that it will be useful
// but WITHOUT ANY WARRANTY; without even the implied warranty of
// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
// GNU Affero General Public License for more details.
//
// You should have received a copy of the GNU Affero General Public License
// along with this program. If not, see <http://www.gnu.org/licenses/>.
package cmd
import (
"os"
"testing"
"time"
"github.com/minio/cli"
xhttp "github.com/minio/minio/internal/http"
)
func TestServerReadHeaderTimeoutConfig(t *testing.T) {
for _, tc := range []struct {
name, env, idle string
args []string
want time.Duration
fmtgen bool
}{
{name: "default", want: xhttp.DefaultReadHeaderTimeout},
{name: "flag", args: []string{"--read-header-timeout=100ms"}, want: 100 * time.Millisecond},
{name: "environment", env: "170ms", want: 170 * time.Millisecond},
{name: "flag-over-environment", env: "170ms", args: []string{"--read-header-timeout=80ms"}, want: 80 * time.Millisecond},
{name: "yaml-retains-flag", args: []string{"--config=testdata/config/1.yaml", "--read-header-timeout=100ms"}, want: 100 * time.Millisecond},
{name: "zero-fallback", args: []string{"--read-header-timeout=0s"}},
{name: "zero-idle-default-header", idle: "0s", want: xhttp.DefaultReadHeaderTimeout},
{name: "negative-idle-default-header", idle: "-1s", want: xhttp.DefaultReadHeaderTimeout},
{name: "fmt-gen-unregistered-duration", env: "100ms", fmtgen: true},
{name: "negative-disabled", args: []string{"--read-header-timeout=-1s"}, want: -time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
for _, key := range []string{"MINIO_ARGS", "MINIO_VOLUMES", "MINIO_ENDPOINTS", "MINIO_CONFIG", "MINIO_ERASURE_SET_DRIVE_COUNT"} {
t.Setenv(key, "")
}
t.Setenv("MINIO_READ_HEADER_TIMEOUT", tc.env)
if tc.env == "" {
if err := os.Unsetenv("MINIO_READ_HEADER_TIMEOUT"); err != nil {
t.Fatal(err)
}
}
idle := tc.idle
if idle == "" {
idle = "2s"
}
t.Setenv("MINIO_IDLE_TIMEOUT", idle)
idleWant, err := time.ParseDuration(idle)
if err != nil {
t.Fatal(err)
}
commandName, flags := "server", serverCmd.Flags
if tc.fmtgen {
commandName, flags = "fmt-gen", fmtGenFlags
idleWant = 0
}
var got serverCtxt
called := false
app := cli.NewApp()
app.Commands = []cli.Command{{Name: commandName, Flags: flags, Action: func(ctx *cli.Context) error {
called = true
if parsed := ctx.Duration("read-header-timeout"); parsed != tc.want {
t.Errorf("CLI read-header-timeout=%s, expected=%s", parsed, tc.want)
}
return buildServerCtxt(ctx, &got)
}}}
args := append([]string{"silo", commandName}, tc.args...)
args = append(args, t.TempDir())
if err := app.Run(args); err != nil {
t.Fatal(err)
}
if !called {
t.Fatal("server action did not run")
}
if got.ReadHeaderTimeout != tc.want {
t.Errorf("parsed ReadHeaderTimeout = %s, want %s", got.ReadHeaderTimeout, tc.want)
}
if got.IdleTimeout != idleWant {
t.Errorf("parsed IdleTimeout = %s, want %s", got.IdleTimeout, idleWant)
}
})
}
}
@@ -274,6 +274,10 @@ func TestBucketMetadataInitialSyncPhysicalCreated(t *testing.T) {
}
events = append(events, event)
}
if r.URL.Path == "/minio/admin/v3/site-replication/peer/iam-revisions" {
_ = json.NewEncoder(w).Encode(iamRevisionResponse{iamRevisionStatus: iamRevisionStatus{Version: iamRevisionProtocol, Node: "initial-peer", Instance: "initial-boot", Digest: "ack"}})
return
}
w.WriteHeader(http.StatusOK)
}))
defer peer.Close()
+87 -34
View File
@@ -28,6 +28,7 @@ import (
"fmt"
"maps"
"math/rand"
"net/http"
"net/url"
"reflect"
"runtime"
@@ -209,7 +210,11 @@ type SiteReplicationSys struct {
// In-memory and persisted multi-site replication state.
state srState
iamMetaCache srIAMCache
iamMetaCache srIAMCache
healOnce sync.Once // Configuration reloads must not spawn more healing loops.
iamHealMu sync.Mutex
iamRevisionProgress map[string]iamRevisionProgress
iamRevisionMetrics iamRevisionMetrics
}
type srState srStateV1
@@ -234,7 +239,7 @@ type srStateData struct {
// Init - initialize the site replication manager.
func (c *SiteReplicationSys) Init(ctx context.Context, objAPI ObjectLayer) error {
go c.startHealRoutine(ctx, objAPI)
c.healOnce.Do(func() { go c.startHealRoutine(ctx, objAPI) })
r := rand.New(rand.NewSource(time.Now().UnixNano()))
for {
err := c.loadFromDisk(ctx, objAPI)
@@ -1257,13 +1262,36 @@ func (c *SiteReplicationSys) IAMChangeHook(ctx context.Context, item madmin.SRIA
return nil
}
if path := iamDeletionPath(item); path != "" {
r, err := loadIAMRevision(ctx, globalIAMSys.store, path)
if err != nil {
return err
}
if !r.Deleted && ((item.Type != madmin.SRIAMItemIAMUser && item.Type != madmin.SRIAMItemGroupInfo) || r.RevokedBefore.IsZero()) {
// A concurrent recreation has already superseded this delete.
return nil
}
item.UpdatedAt = r.timestamp()
if !r.Deleted {
item.UpdatedAt = r.RevokedBefore
}
}
versioned, err := c.replicationItem(ctx, item)
if errors.Is(err, errIAMStaleUpdate) {
return nil
}
if err != nil {
return err
}
cerr := c.concDo(nil, func(d string, p madmin.PeerInfo) error {
admClient, err := c.getAdminClient(ctx, d)
if err != nil {
return wrapSRErr(err)
}
return c.annotatePeerErr(p.Name, replicateIAMItem, admClient.SRPeerReplicateIAMItem(ctx, item))
_, err = executeIAMRevisionRequest(ctx, admClient, http.MethodPut, &iamRevisionBatch{Version: iamRevisionProtocol, Items: []iamReplicationItem{versioned}})
return c.annotatePeerErr(p.Name, replicateIAMItem, err)
},
replicateIAMItem,
)
@@ -1273,6 +1301,7 @@ func (c *SiteReplicationSys) IAMChangeHook(ctx context.Context, item madmin.SRIA
// PeerAddPolicyHandler - copies IAM policy to local. A nil policy argument,
// causes the named policy to be deleted.
func (c *SiteReplicationSys) PeerAddPolicyHandler(ctx context.Context, policyName string, p *policy.Policy, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
var err error
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
@@ -1286,18 +1315,19 @@ func (c *SiteReplicationSys) PeerAddPolicyHandler(ctx context.Context, policyNam
_, err = globalIAMSys.SetPolicy(ctx, policyName, *p)
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerIAMUserChangeHandler - copies IAM user to local.
func (c *SiteReplicationSys) PeerIAMUserChangeHandler(ctx context.Context, change *madmin.SRIAMUser, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
if !change.IsDeleteReq && !updatedAt.IsZero() {
if ui, err := globalIAMSys.GetUserInfo(ctx, change.AccessKey); err == nil && ui.UpdatedAt.After(updatedAt) {
return nil
}
@@ -1326,13 +1356,14 @@ func (c *SiteReplicationSys) PeerIAMUserChangeHandler(ctx context.Context, chang
}
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerGroupInfoChangeHandler - copies group changes to local.
func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, change *madmin.SRGroupInfo, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1340,7 +1371,7 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
var err error
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
if !updatedAt.IsZero() && (!updReq.IsRemove || len(updReq.Members) != 0) {
if gd, err := globalIAMSys.GetGroupDescription(updReq.Group); err == nil && gd.UpdatedAt.After(updatedAt) {
return nil
}
@@ -1349,7 +1380,8 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
if updReq.IsRemove {
_, err = globalIAMSys.RemoveUsersFromGroup(ctx, updReq.Group, updReq.Members)
} else {
if updReq.Status != "" && len(updReq.Members) == 0 {
snapshot, _ := ctx.Value(iamGroupSnapshotKey{}).(bool)
if !snapshot && updReq.Status != "" && len(updReq.Members) == 0 {
_, err = globalIAMSys.SetGroupStatus(ctx, updReq.Group, updReq.Status == madmin.GroupEnabled)
} else {
if globalIAMSys.LDAPConfig.Enabled() {
@@ -1365,13 +1397,14 @@ func (c *SiteReplicationSys) PeerGroupInfoChangeHandler(ctx context.Context, cha
}
}
if err != nil && !errors.Is(err, errNoSuchGroup) {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerSvcAccChangeHandler - copies service-account change to local.
func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change *madmin.SRSvcAccChange, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if change == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1382,7 +1415,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
if len(change.Create.SessionPolicy) > 0 {
sp, err = policy.ParseConfig(bytes.NewReader(change.Create.SessionPolicy))
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
// skip overwrite of local update if peer sent stale info
@@ -1394,6 +1427,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
opts := newServiceAccountOpts{
accessKey: change.Create.AccessKey,
secretKey: change.Create.SecretKey,
status: change.Create.Status,
sessionPolicy: sp,
claims: change.Create.Claims,
name: change.Create.Name,
@@ -1402,7 +1436,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
}
_, _, err = globalIAMSys.NewServiceAccount(ctx, change.Create.Parent, change.Create.Groups, opts)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
case change.Update != nil:
@@ -1411,7 +1445,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
if len(change.Update.SessionPolicy) > 0 {
sp, err = policy.ParseConfig(bytes.NewReader(change.Update.SessionPolicy))
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
// skip overwrite of local update if peer sent stale info
@@ -1431,7 +1465,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
_, err = globalIAMSys.UpdateServiceAccount(ctx, change.Update.AccessKey, opts)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
case change.Delete != nil:
@@ -1442,7 +1476,7 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
}
}
if err := globalIAMSys.DeleteServiceAccount(ctx, change.Delete.AccessKey, true); err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
@@ -1451,24 +1485,17 @@ func (c *SiteReplicationSys) PeerSvcAccChangeHandler(ctx context.Context, change
// PeerPolicyMappingHandler - copies policy mapping to local.
func (c *SiteReplicationSys) PeerPolicyMappingHandler(ctx context.Context, mapping *madmin.SRPolicyMapping, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if mapping == nil {
return errSRInvalidRequest(errInvalidArgument)
}
// skip overwrite of local update if peer sent stale info
if !updatedAt.IsZero() {
mp, ok := globalIAMSys.store.GetMappedPolicy(mapping.Policy, mapping.IsGroup)
if ok && mp.UpdatedAt.After(updatedAt) {
return nil
}
}
// When LDAP is enabled, we verify that the user or group exists in LDAP and
// use the normalized form of the entityName (which will be an LDAP DN).
userType := IAMUserType(mapping.UserType)
isGroup := mapping.IsGroup
entityName := mapping.UserOrGroup
if globalIAMSys.GetUsersSysType() == LDAPUsersSysType && userType == stsUser {
if mapping.Policy != "" && globalIAMSys.GetUsersSysType() == LDAPUsersSysType && userType == stsUser {
// Validate that the user or group exists in LDAP and use the normalized
// form of the entityName (which will be an LDAP DN).
var err error
@@ -1491,19 +1518,20 @@ func (c *SiteReplicationSys) PeerPolicyMappingHandler(ctx context.Context, mappi
entityName = foundUserDN.NormDN
}
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
}
_, err := globalIAMSys.PolicyDBSet(ctx, entityName, mapping.Policy, userType, isGroup)
if err != nil {
return wrapSRErr(err)
return iamReplicationError(err)
}
return nil
}
// PeerSTSAccHandler - replicates STS credential locally.
func (c *SiteReplicationSys) PeerSTSAccHandler(ctx context.Context, stsCred *madmin.SRSTSCredential, updatedAt time.Time) error {
ctx = withIAMReplicationTime(ctx, updatedAt)
if stsCred == nil {
return errSRInvalidRequest(errInvalidArgument)
}
@@ -1555,7 +1583,7 @@ func (c *SiteReplicationSys) PeerSTSAccHandler(ctx context.Context, stsCred *mad
// Set these credentials to IAM.
if _, err := globalIAMSys.SetTempUser(ctx, cred.AccessKey, cred, stsCred.ParentPolicyMapping); err != nil {
return fmt.Errorf("unable to save STS credential and/or parent policy mapping: %w", err)
return iamReplicationError(fmt.Errorf("unable to save STS credential and/or parent policy mapping: %w", err))
}
return nil
@@ -4193,7 +4221,8 @@ func (c *SiteReplicationSys) SiteReplicationMetaInfo(ctx context.Context, objAPI
}
info.UserInfoMap[k] = madmin.UserInfo{
Status: madmin.AccountStatus(v.Credentials.Status),
Status: madmin.AccountStatus(v.Credentials.Status),
UpdatedAt: v.UpdatedAt,
}
}
}
@@ -4549,9 +4578,25 @@ func (c *SiteReplicationSys) PeerStateEditReq(ctx context.Context, arg madmin.SR
const siteHealTimeInterval = 30 * time.Second
func (c *SiteReplicationSys) startHealRoutine(ctx context.Context, objAPI ObjectLayer) {
ctx, cancel := globalLeaderLock.GetLock(ctx)
defer cancel()
for ctx.Err() == nil {
var leadership LockContext
select {
case <-ctx.Done():
return
case leadership = <-globalLeaderLock.lockContext:
}
if leadership.Context().Err() != nil {
continue
}
leaderCtx, cancel := mergeContext(leadership.Context(), ctx)
c.healWithLeadership(leaderCtx, objAPI)
cancel()
// Quorum loss cancels a leadership lease, not the subsystem. Wait
// for a new lease so revocations can still reach offline sites.
}
}
func (c *SiteReplicationSys) healWithLeadership(ctx context.Context, objAPI ObjectLayer) {
healTimer := time.NewTimer(siteHealTimeInterval)
defer healTimer.Stop()
@@ -5170,13 +5215,16 @@ func (c *SiteReplicationSys) healBucketReplicationConfig(ctx context.Context, ob
}
func (c *SiteReplicationSys) healIAMSystem(ctx context.Context, objAPI ObjectLayer) error {
// A peer rejecting a deletion must not stop unrelated live updates from
// reaching healthy peers. Retain and report the error for the next retry.
deletionErr := c.healIAMDeletions(ctx)
info, err := c.siteReplicationStatus(ctx, objAPI, madmin.SRStatusOptions{
Users: true,
Policies: true,
Groups: true,
})
if err != nil {
return err
return errors.Join(deletionErr, err)
}
for policy := range info.PolicyStats {
c.healPolicies(ctx, objAPI, policy, info)
@@ -5194,7 +5242,7 @@ func (c *SiteReplicationSys) healIAMSystem(ctx context.Context, objAPI ObjectLay
c.healGroupPolicies(ctx, objAPI, group, info)
}
return nil
return deletionErr
}
// heal iam policies present on this site to peers, provided current cluster has the most recent update.
@@ -5429,8 +5477,10 @@ func (c *SiteReplicationSys) healUsers(ctx context.Context, objAPI ObjectLayer,
peerName := info.Sites[dID].Name
u, ok := globalIAMSys.GetUser(ctx, user)
if !ok {
// Disabled identities are valid replication sources. CheckKey returns
// their stored record even though authentication is denied.
u, _, err := globalIAMSys.CheckKey(ctx, user)
if err != nil || u.Credentials.AccessKey == "" || u.Credentials.IsExpired() {
continue
}
creds := u.Credentials
@@ -5629,7 +5679,10 @@ func isGroupDescEqual(g1, g2 madmin.GroupDesc) bool {
}
func isUserInfoEqual(u1, u2 madmin.UserInfo) bool {
if u1.PolicyName != u2.PolicyName ||
// Full-site summaries omit secrets and claims. Equal status alone cannot
// distinguish a recreated identity or an edited service-account policy.
if !u1.UpdatedAt.Equal(u2.UpdatedAt) ||
u1.PolicyName != u2.PolicyName ||
u1.Status != u2.Status ||
u1.SecretKey != u2.SecretKey {
return false
+9
View File
@@ -109,6 +109,15 @@ func testStorageAPIListDir(t *testing.T, storage StorageAPI) {
}
}
}
for _, name := range []string{"one", "two", "three"} {
if err := storage.AppendFile(t.Context(), "foo", "bounded/"+name, []byte("x")); err != nil {
t.Fatal(err)
}
}
entries, err := storage.ListDir(t.Context(), "", "foo", "bounded", 2)
if err != nil || len(entries) != 2 {
t.Fatalf("ListDir count was lost in storage/RPC path: %v %v", entries, err)
}
}
func testStorageAPIReadAll(t *testing.T, storage StorageAPI) {
+4
View File
@@ -624,6 +624,10 @@ func (sts *stsAPIHandlers) AssumeRole(w http.ResponseWriter, r *http.Request) {
claims[expClaim] = UTCNow().Add(duration).Unix()
claims[parentClaim] = user.AccessKey
if err := setIAMParentRevocationClaim(ctx, globalIAMSys.store, user.AccessKey, claims); err != nil {
writeSTSErrorResponse(ctx, w, ErrSTSInternalError, err)
return
}
tokenRevokeType := r.Form.Get(stsRevokeTokenType)
if tokenRevokeType != "" {
@@ -2,7 +2,7 @@
PR #60 introduced an opt-in scheduler that moved objects between local server pools according to GET frequency. It and its feature-specific fixes have been removed. This does not remove ordinary lifecycle expiration, transitions to remote tiers, rebalance, decommission, or the general multi-pool correctness fixes from PR #178.
The [introduction and rollback record](../../investigations/access-tiering-revert.md) documents the commit history, scope decision, review corrections and unresolved validation findings.
The [introduction and rollback record](https://silo.pgsty.com/compatibility/access-tiering-removal/) documents the commit history, scope decision, review corrections and unresolved validation findings.
The published Server 20260903 predates this feature. These instructions concern main/snapshot deployments that included #60; upgrading from the published version does not require access-tier configuration cleanup.
@@ -35,7 +35,7 @@ set -e
export MINIO_CI_CD=1
export MINIO_BROWSER=off
export MINIO_ROOT_USER="minio"
export MINIO_ROOT_PASSWORD="silo123"
export MINIO_ROOT_PASSWORD="silo12345"
export MINIO_KMS_AUTO_ENCRYPTION=off
export MINIO_PROMETHEUS_AUTH_TYPE=public
export MINIO_KMS_SECRET_KEY=my-minio-key:OSMM+vkKUTCvQs9YL/CVMIMt43HFhkUpqJxTmGl6rYw=
@@ -58,8 +58,8 @@ silo server --address 127.0.0.1:9003 "http://127.0.0.1:9003/tmp/multisiteb/data/
silo server --address 127.0.0.1:9004 "http://127.0.0.1:9003/tmp/multisiteb/data/disterasure/xl{1...4}" \
"http://127.0.0.1:9004/tmp/multisiteb/data/disterasure/xl{5...8}" >/tmp/siteb_2.log 2>&1 &
export MC_HOST_sitea=http://minio:silo123@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo123@127.0.0.1:9004
export MC_HOST_sitea=http://minio:silo12345@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo12345@127.0.0.1:9004
./mc ready sitea
./mc ready siteb
@@ -83,7 +83,7 @@ done
echo "adding replication rule for site a -> site b"
./mc replicate add sitea/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/bucket
--remote-bucket http://minio:silo12345@127.0.0.1:9004/bucket
remote_arn=$(./mc replicate ls sitea/bucket --json | jq -r .rule.Destination.Bucket)
sleep 1
@@ -137,7 +137,7 @@ fi
echo "adding replication rule for site a -> site b"
./mc replicate add sitea/bucket-version/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/bucket-version
--remote-bucket http://minio:silo12345@127.0.0.1:9004/bucket-version
./mc mb sitea/bucket-version/directory/
@@ -37,7 +37,7 @@ set -e
export MINIO_CI_CD=1
export MINIO_BROWSER=off
export MINIO_ROOT_USER="minio"
export MINIO_ROOT_PASSWORD="silo123"
export MINIO_ROOT_PASSWORD="silo12345"
export MINIO_KMS_AUTO_ENCRYPTION=off
export MINIO_PROMETHEUS_AUTH_TYPE=public
export MINIO_KMS_SECRET_KEY=my-minio-key:OSMM+vkKUTCvQs9YL/CVMIMt43HFhkUpqJxTmGl6rYw=
@@ -70,9 +70,9 @@ silo server --address 127.0.0.1:9005 "http://127.0.0.1:9005/tmp/multisitec/data/
silo server --address 127.0.0.1:9006 "http://127.0.0.1:9005/tmp/multisitec/data/disterasure/xl{1...4}" \
"http://127.0.0.1:9006/tmp/multisitec/data/disterasure/xl{5...8}" >/tmp/sitec_2.log 2>&1 &
export MC_HOST_sitea=http://minio:silo123@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo123@127.0.0.1:9004
export MC_HOST_sitec=http://minio:silo123@127.0.0.1:9006
export MC_HOST_sitea=http://minio:silo12345@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo12345@127.0.0.1:9004
export MC_HOST_sitec=http://minio:silo12345@127.0.0.1:9006
./mc ready sitea
./mc ready siteb
@@ -93,73 +93,73 @@ export MC_HOST_sitec=http://minio:silo123@127.0.0.1:9006
echo "adding replication rule for a -> b : ${remote_arn}"
sleep 1
./mc replicate add sitea/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9004/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync"
sleep 1
echo "adding replication rule for b -> a : ${remote_arn}"
./mc replicate add siteb/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9001/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9001/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync"
sleep 1
echo "adding replication rule for a -> c : ${remote_arn}"
./mc replicate add sitea/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9006/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9006/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 2
sleep 1
echo "adding replication rule for c -> a : ${remote_arn}"
./mc replicate add sitec/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9001/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9001/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 2
sleep 1
echo "adding replication rule for b -> c : ${remote_arn}"
./mc replicate add siteb/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9006/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9006/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 3
sleep 1
echo "adding replication rule for c -> b : ${remote_arn}"
./mc replicate add sitec/bucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/bucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9004/bucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 3
sleep 1
echo "adding replication rule for olockbucket a -> b : ${remote_arn}"
./mc replicate add sitea/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9004/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync"
sleep 1
echo "adding replication rule for olockbucket b -> a : ${remote_arn}"
./mc replicate add siteb/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9001/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9001/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync"
sleep 1
echo "adding replication rule for olockbucket a -> c : ${remote_arn}"
./mc replicate add sitea/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9006/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9006/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 2
sleep 1
echo "adding replication rule for olockbucket c -> a : ${remote_arn}"
./mc replicate add sitec/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9001/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9001/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 2
sleep 1
echo "adding replication rule for olockbucket b -> c : ${remote_arn}"
./mc replicate add siteb/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9006/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9006/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 3
sleep 1
echo "adding replication rule for olockbucket c -> b : ${remote_arn}"
./mc replicate add sitec/olockbucket/ \
--remote-bucket http://minio:silo123@127.0.0.1:9004/olockbucket \
--remote-bucket http://minio:silo12345@127.0.0.1:9004/olockbucket \
--replicate "existing-objects,delete,delete-marker,replica-metadata-sync" --priority 3
sleep 1
@@ -204,33 +204,33 @@ head -c 221227088 </dev/urandom >200M
sleep 10
echo "Verifying ETag for all objects"
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9001/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9002/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9003/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9004/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9005/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9006/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9001/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9002/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9003/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9004/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9005/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9006/ -bucket bucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9001/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9002/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9003/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9004/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9005/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo123 -endpoint http://127.0.0.1:9006/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9001/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9002/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9003/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9004/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9005/ -bucket olockbucket
./s3-check-md5 -versions -access-key minio -secret-key silo12345 -endpoint http://127.0.0.1:9006/ -bucket olockbucket
# additional tests for encryption object alignment
go install -v github.com/minio/multipart-debug@latest
upload_id=$(multipart-debug --endpoint 127.0.0.1:9001 --accesskey minio --secretkey silo123 multipart new --bucket bucket --object new-test-encrypted-object --encrypt)
upload_id=$(multipart-debug --endpoint 127.0.0.1:9001 --accesskey minio --secretkey silo12345 multipart new --bucket bucket --object new-test-encrypted-object --encrypt)
dd if=/dev/urandom bs=1 count=7048531 of=/tmp/7048531.txt
dd if=/dev/urandom bs=1 count=2847391 of=/tmp/2847391.txt
sudo apt install jq -y
etag_1=$(multipart-debug --endpoint 127.0.0.1:9002 --accesskey minio --secretkey silo123 multipart upload --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} --file /tmp/7048531.txt --number 1 | jq -r .ETag)
etag_2=$(multipart-debug --endpoint 127.0.0.1:9001 --accesskey minio --secretkey silo123 multipart upload --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} --file /tmp/2847391.txt --number 2 | jq -r .ETag)
multipart-debug --endpoint 127.0.0.1:9002 --accesskey minio --secretkey silo123 multipart complete --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} 1.${etag_1} 2.${etag_2}
etag_1=$(multipart-debug --endpoint 127.0.0.1:9002 --accesskey minio --secretkey silo12345 multipart upload --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} --file /tmp/7048531.txt --number 1 | jq -r .ETag)
etag_2=$(multipart-debug --endpoint 127.0.0.1:9001 --accesskey minio --secretkey silo12345 multipart upload --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} --file /tmp/2847391.txt --number 2 | jq -r .ETag)
multipart-debug --endpoint 127.0.0.1:9002 --accesskey minio --secretkey silo12345 multipart complete --bucket bucket --object new-test-encrypted-object --uploadid ${upload_id} 1.${etag_1} 2.${etag_2}
sleep 10
@@ -35,7 +35,7 @@ set -e
export MINIO_CI_CD=1
export MINIO_BROWSER=off
export MINIO_ROOT_USER="minio"
export MINIO_ROOT_PASSWORD="silo123"
export MINIO_ROOT_PASSWORD="silo12345"
export MINIO_KMS_AUTO_ENCRYPTION=off
export MINIO_PROMETHEUS_AUTH_TYPE=public
export MINIO_KMS_SECRET_KEY=my-minio-key:OSMM+vkKUTCvQs9YL/CVMIMt43HFhkUpqJxTmGl6rYw=
@@ -70,10 +70,10 @@ silo server --address 127.0.0.1:9008 "http://127.0.0.1:9007/tmp/multisited/data/
# Wait to make sure all Silo instances are up
export MC_HOST_sitea=http://minio:silo123@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo123@127.0.0.1:9004
export MC_HOST_sitec=http://minio:silo123@127.0.0.1:9006
export MC_HOST_sited=http://minio:silo123@127.0.0.1:9008
export MC_HOST_sitea=http://minio:silo12345@127.0.0.1:9001
export MC_HOST_siteb=http://minio:silo12345@127.0.0.1:9004
export MC_HOST_sitec=http://minio:silo12345@127.0.0.1:9006
export MC_HOST_sited=http://minio:silo12345@127.0.0.1:9008
./mc ready sitea
./mc ready siteb
@@ -89,7 +89,7 @@ export MC_HOST_sited=http://minio:silo123@127.0.0.1:9008
sleep 10s
## Add warm tier
./mc ilm tier add minio sitea WARM-TIER --endpoint http://localhost:9006 --access-key minio --secret-key silo123 --bucket bucket
./mc ilm tier add minio sitea WARM-TIER --endpoint http://localhost:9006 --access-key minio --secret-key silo12345 --bucket bucket
## Add ILM rules
./mc ilm add sitea/bucket --transition-days 0 --transition-tier WARM-TIER --transition-days 0 --noncurrent-expire-days 2 --expire-days 3 --prefix "myprefix" --tags "tag1=val1&tag2=val2"
+2
View File
@@ -1,5 +1,7 @@
# Compression Guide
For SILO's current SSE-C compression and historical-object boundary, see the [maintained encryption guide](https://silo.pgsty.com/administration/server-side-encryption/server-side-encryption-sse-c/) and [SSE-C replica design](https://silo.pgsty.com/blog/design/ssec-replica-integrity/). Check the documented release boundary before applying main-branch behavior to an older binary.
Silo server allows streaming compression to ensure efficient disk space usage.
Compression happens inflight, i.e objects are compressed before being written to disk(s).
Silo uses [`klauspost/compress/s2`](https://github.com/klauspost/compress/tree/master/s2)
+2
View File
@@ -1,5 +1,7 @@
# Federation Quickstart Guide *Federation feature is deprecated and should be avoided for future deployments*
The maintained [federated CopyObject design](https://silo.pgsty.com/blog/design/federated-copy-object/) records the destination encryption, checksum, Object Lock and committed-response contract, including its Server release boundary.
This document explains how to configure Silo with `Bucket lookup from DNS` style federation.
## Cross-deployment copy behavior
+2 -2
View File
@@ -38,7 +38,7 @@ pid=$!
mc ready mysilo
mc admin user add mysilo/ silo123 silo123
mc admin user add mysilo/ silo123 silo12345
mc admin policy create mysilo/ deny-non-sse-kms-pol ./docs/iam/policies/deny-non-sse-kms-objects.json
mc admin policy create mysilo/ deny-invalid-sse-kms-pol ./docs/iam/policies/deny-objects-with-invalid-sse-kms-key-id.json
@@ -50,7 +50,7 @@ mc admin policy attach mysilo consoleAdmin --user silo123
mc mb -l mysilo/test-bucket
mc mb -l mysilo/multi-key-poc
export MC_HOST_mysilo1="http://silo123:silo123@localhost:9000/"
export MC_HOST_mysilo1="http://silo123:silo12345@localhost:9000/"
mc cp /etc/issue mysilo1/test-bucket
ret=$?
@@ -1,193 +0,0 @@
# 访问频率分层:引入、修复与回退记录
记录日期:2026-09-15。本记录说明 PR #60 的设计、合入后的取舍,以及此次回退为什么同时保留并补齐通用多池正确性修复。操作步骤见[退役迁移说明](../bucket/lifecycle/access-tiering-removal.md)。合并、正式发布和生产部署是不同状态;本记录随回退变更交付,不代表已经发布。
初稿审查时(2026-09-15 03:20 UTC),最终候选尚未移植到远端基线、尚未冻结 PR head,合并前全量检查与三次有效 Linux 运行尚未执行,本变更尚未合并。后续执行状态以承载本记录的 PR 及其绑定提交的验收记录为准;下文的历史实验不替代这些检查。
**当前验收状态:** 回退、普通版本 DELETE 调和、扫描复用及文档由 [PR #188](https://github.com/pgsty/silo/pull/188) 交付。本地与 CI 检查通过;首次 Linux 验收因 DELETE204 后的 HEAD/GET503 停止。随后完成可控机制实验、匹配写入负载对照,并修正准备检查;独立的新轮次 R-Upgrade-2 三次完整升级验收均通过。旧失败没有改判,原单次请求的逐盘状态仍不可追溯。最终合并状态以 PR 为准,正式发布与部署另行验收。详情见第 9 至 11 节。
## 1. 引入的目标和实际范围
[@mrjavadseydi](https://github.com/mrjavadseydi) 在 [PR #60](https://github.com/pgsty/silo/pull/60) 提出了基于 GET 频率的本地池间分层。成功 GET 更新有界滚动计数,后台调度器把热对象提升到配置中的首个池,把已经迁移且变冷的对象降到末个池。功能默认关闭,至少需要两个池;它与普通生命周期过期、远端对象存储 transition、rebalance 和 decommission 是不同机制。
实现不只是一个后台任务:它增加了 GET 计数入口、leader 调度、跨池版本栈复制、来源清理及失败恢复、热池配额、生命周期 XML 扩展、配置项、指标和扫描统计。访问热度统计使 data-usage cache 从 v8 升到 v9。对象本体的存储格式没有因此改变。
搬移要保持完整版本历史、删除标记、null version、时间戳、ETag、校验和和加密元数据;同时需要处理并发写入、目的端已有版本、来源部分删除和远端 tier 引用。合入前的修补和测试针对并覆盖了这些边界,不应把回退解释成贡献无效。贡献者署名继续保留,独立的其他贡献也不回退。
## 2. 可追溯时间线
下表的 PR/Issue 时间使用 UTC;提交链接对应具体代码,不把“报告时间”当作“缺陷首次出现时间”。
| 时间 | 事件 | 本次处理 |
| --- | --- | --- |
| 2026-08-15 10:15 | #60 创建;原始实现 [`7a060cab1`](https://github.com/pgsty/silo/commit/7a060cab1edd5bbc17da7f703bbd1ab7415b6f7c) | 随特性撤销 |
| 2026-09-06 00:09 | [#133](https://github.com/pgsty/silo/issues/133) 报告多池副本写不能权威调和 Object Lock 状态 | 保留解决它的通用修复 |
| 2026-09-06 15:12 | [#144](https://github.com/pgsty/silo/issues/144) 报告条件 DELETE 原子性仅限单个纠删码集合 | 保留解决它的通用修复 |
| 2026-09-08 05:18 | [`9a6e1477f`](https://github.com/pgsty/silo/commit/9a6e1477f45067559def8423d431ee177795134f) 补访问分层兼容标识清单 | 删除功能专属标识,保留有依据的退役兼容 |
| 2026-09-08 07:08 | [`374de0fa3`](https://github.com/pgsty/silo/commit/374de0fa32aa1d6eda57dbba6ac522ebf793b6be) 修复搬移的版本保全、写隔离和删除范围 | 随专属搬移器撤销 |
| 2026-09-08 07:28 | [`5ac33e158`](https://github.com/pgsty/silo/commit/5ac33e1583e838ade3f56c7d60e80e7a854f9a88) 确定性覆盖搬移失败恢复 | 随已删除搬移器的专属测试撤销 |
| 2026-09-08 07:42 | #60 以 [`a3df317ae`](https://github.com/pgsty/silo/commit/a3df317ae0725eb650e4d3e21551154f69be6229) 合入 | 以该 merge 的第一父差异确定功能边界 |
| 2026-09-11 12:34 | [#178](https://github.com/pgsty/silo/pull/178) 合入通用多池写入、元数据与条件删除调和 | 保留,解除测试对访问搬移器的依赖 |
| 2026-09-13 | [`2dd1e00da`](https://github.com/pgsty/silo/commit/2dd1e00da49faf995f2db29807fc6211b8376d7d) 的 CHANGELOG 同时汇总访问分层和通用多池修复 | 拆开表述,不整条删除独立修复历史 |
| 2026-09-15 | 维护者决定收缩访问分层;完成来源分析、三种候选反证、退役兼容、普通版本 DELETE 修补与外部评审 | 形成此次选择性回退 |
#178 中的 [`e59a3d938`](https://github.com/pgsty/silo/commit/e59a3d938ed25c1bcd51efbb4ad6955073d195f7) 提供池级串行化、字段调和和条件删除;[`ccb676e60`](https://github.com/pgsty/silo/commit/ccb676e60cb7441ee65ff7c35f3b7828979101fd) 保护仍被其他副本使用的远端 tier 引用;[`51d41345f`](https://github.com/pgsty/silo/commit/51d41345f7ac532f9c6fea2b7dba5f29da930b8f) 保存 Linux 重启及 OIDC 验收记录。这三项不属于仅为访问频率调度而存在的代码。
截至回退评估时,公开 Server `RELEASE.2026-09-03T13-18-01Z` 早于 #60 合入。退役迁移主要针对运行过后续 main、自行构建或快照版本的实例;不能据此声称正式 release 用户普遍启用过此特性。
## 3. 为什么回退,为什么不能整批撤销后续修复
维护者的取舍是:默认关闭的可选调度能力,对配置、生命周期、缓存、统计和核心多池写入路径带来了过大的维护面。此次移除的是该能力及专属实现,没有测量并宣称吞吐量提升、延迟降低或固定减少一次分布式锁往返。
“后来改过同一文件”不等于“由 #60 引发”。#133/#144 的报告早于 #60 合入;更关键的是,原有 rebalance、decommission 和复制写入也会使同一版本暂时存在于多个池。访问分层消失后,这些状态仍然合法存在。移除调和与锁纪律会重新允许旧副本遮蔽新元数据、条件删除选错版本、清理错误被吞掉等问题。
评估在隔离工作树中实际比较了三条路线:
| 候选 | 通用多池回归 | 普通指定版本 DELETE |
| --- | --- | --- |
| A:撤销 #60,保留 #178 | 原有 13 组通过 | 仍能成功返回后留下可读副本 |
| B:同时撤销 #60 和 #178 的存储修补 | 相同 13 组中 10 组失败 | 问题仍在 |
| C:A 加普通版本 DELETE 调和 | 13 组原有及当时新增的 8 组通过 | 同一复现通过 |
这些是 2026-09-15 的历史对照结果,不是最终 PR head 的发布验收。后续补上目录标记、真实 rebalance 中断等覆盖后,通用多池测试达到 23 组。测试通过不能替代来源分析,来源分析也不能替代最终候选的运行验证。
## 4. 最终保留与删除的边界
- 删除访问 tracker、调度/搬移器、GET 和 scanner 钩子、热池配额、专属配置帮助、生命周期动作、指标及专属测试。
- 保留普通过期、远端 transition、rebalance/decommission、复制写入,以及 #178 的 Object Lock、标签、条件删除、元数据调和与远端引用保护。
- 原有 data-usage 生成解码器恢复到 #60 前的实现;允许读取 v8/v9,利用字段编码跳过已退役热度字段,继续写 v8。普通字段保真由历史真实 v9 样本测试覆盖。
- 仅容忍准确的十个退役 ILM 键;读取生命周期时丢弃退役扩展。纯访问动作的规则需先清理才能再次编辑;混合规则保留普通动作。
- 已经搬移的对象留在当前池;没有全量搬回、自动删除所有重复版本、后台清理服务或对象元数据重写。
曾对本地审查提交 `d06f1c614` 做过声明来源复核:消失的 357 个声明均不在 #60 之前,删除的十个文件均由 #60 引入;#60 原先删换的 41 行旧文本按忽略空白比较有 40 行恢复,剩下一行保留 #178 在已持有池锁时调用 `getWritePoolIdx(..., true)` 的修正,避免对同一对象再次取锁。生成缓存解码器与功能前逐字节一致,原有 13 组通用测试没有删除。这是该审查版本的保全证据,不能把数字脱离 SHA 当作未来所有版本的保证。
## 5. 普通版本 DELETE 是独立补洞
功能移除不会自动消除历史重复副本。普通单对象指定版本 DELETE 因而复用既有调和路径:在池级对象锁内读取每个池的目标版本,计算一次条件及回调,先删除非权威副本,再处理权威副本;任何不可读池或清理错误都不能当作成功。
范围包括 UUID、null version、delete marker,以及原先就被解析为 null version 的未指定版本目录标记 DELETE。入站复制、搬移内部调用、生命周期过期和 free-version 清理保留各自语义;批量 `DeleteObjects` 原本就会向池并发扇出,不是此次遗漏。
删除标记需要向 retention/metadata 回调传入与 set 层相同的 `MethodNotAllowed` 或 `ObjectNotFound` 语义。直接复用拒绝 marker 的元数据更新入口会错误地拒绝合法版本删除。回调从所有副本合并独立更新的 Object Lock 和标签,不能随意只采用一个池的状态。
存在两项明确的成功/失败边界:
1. 读法定多数不足时返回 `503 SlowDownRead`,即使另一个池有可读副本。旧路径的结果会受池遍历顺序影响;新路径把失败语义统一。它是正确性与可用性的取舍,需要恢复后重试。
2. 出站删除复制尚未完成时,成功响应可以表示各副本进入 `VersionPurgePending`,由既有 worker 完成清理。原有每池 quorum 规则也继续适用;不能把成功响应等同于每一块盘立即物理删除。
## 6. 审查如何改变了方案
本地先形成三个线性审查提交:`8fdfdabd9` 移除特性,`d06f1c614` 补普通 DELETE,`6e3fdca97` 去除重复扫描。它们记录审查演进,最终 PR 在独立远端基线上重放,提交 ID 会改变;不应把线性演进误认为三份同时维护的实现。
Claude Code Opus 5 / max 的五轮实现评审要求补齐退役兼容、说明协调停机及环境一致性、验证真实池故障,并纠正运行证据措辞。随后 Claude 与 ZCode 的独立复核再次确认了回退边界和普通 DELETE 语义;最终两轮计划商榷收束了交付流程。
| 意见 | 裁定与处理 |
| --- | --- |
| 指定版本 DELETE 重复扫描所有池 | 接受。第一次扫描已在同一锁内得到目标副本,直接合并其结果。16 盘池在回调前的读取计数从 32 降为 16;这不等于总 I/O 或延迟减半。 |
| N 个副本产生 N 条 DELETE 审计 | 反驳。底层调用只追加上下文标签,HTTP 层向每个配置审计目标发送一次请求事件。成功完成调和时池标签最终指向 primary;不增加额外 NoAuditLog 修改。 |
| retention/metadata 顺序与 set 层不同 | 差异存在,但已有明确注释。保留 retention 优先,拒绝后不删除、不调度删除复制、不 Sweep;不宣称任意自定义回调都能交换。 |
| 把优化 amend 进旧提交并直接丢弃 main 脏修改 | 不 amend 已审历史。先保存完整文件、补丁、哈希及可达恢复引用,核对覆盖后受控恢复。八组新增通用测试承接,访问搬移专用测试由真实 rebalance 中断覆盖替代。 |
| 从当前本地分支直接开 PR | 调整。其祖先含另一个任务的 IAM/超时提交 `ebc9937d9`;从远端 `89637554d` 仅移植本次变更,不夹带或删除独立工作。 |
| 合并之后再跑全量与 Linux 验收 | 不接受。先完成文档和移植,冻结 PR head,再验收;不能把旧 SHA 的测试直接提升为新基线的通过记录。 |
| 不同基线 diff 必须逐字节一致 | 改为每提交 stable patch-id、路径与完整树等价性核验。blob hash 和行号随基线变化,不应成为错误的拒绝依据。 |
`objectPoolInfos` 的并行查询作为独立性能跟进,本次不增加并发实现。Contributor 署名、#132 配额指标、#77 桶元数据、federated COPY、IAM、超时及其他独立修复不因本次取舍被整体撤销。
## 7. 历史运行证据与未定案事项
以下是移植前 `d06f1c614` 的历史实验,不是最终 PR head 的验收替身。
| 运行 | 实际观察 | 不应推导的结论 |
| --- | --- | --- |
| `f590538f`,四节点双池 | 旧版真实 rebalance 中断留下 9 个重复 UUID;协调停机换新后对象四节点可读,旧配置与普通规则可编辑,真实缓存头 v9→v8;删除一个 addressed UUID 后四节点 HEAD/GET 404,另一个版本仍可读 | 没有验证全部 9 个不同 UUID 收敛;物理元数据仅每池抽查一盘;没有整份运行时统计守恒测量 |
| `fd93cd37`,三节点双池 | 停掉一个池所在节点,保留 namespace 锁的 2/3 法定多数;DELETE 返回 503 SlowDownRead,源全部四盘保留目标版本;恢复后 DELETE204,各节点 HEAD/GET404 | 不能推广为跨主机网络、持久盘及压力验收 |
| 被弃用的四节点停池拓扑 | 同时丢失 namespace 锁法定多数,发生客户端超时,记录脚本还遇到 NoneType 错误 | 不是“池读取返回503”的证明 |
| `83676ca2` | 升级后配置/生命周期编辑之后一次 HeadObject 返回503;artifact 没有记录 DELETE 自身响应 | “DELETE之后”仅来自脚本顺序,不能写成已证实 DELETE204 后异常,也不能归类为已修复、既有问题或暂态 |
**开放项:`83676ca2` 的 HEAD503 仍未定案。** 可直接比较的目标阶段历史运行是一失败、一成功;一次未复现不足以关闭问题。仅凭 `SlowDownWrite` 等错误名字也不能给其他失败确定容量或环境根因。
最终候选的合并前复核采用有界规则:要求三次有效升级后 DELETE/HEAD 运行,最多五次总尝试;每次记录 DELETE 码/耗时、失败 HEAD 的节点与版本、GET 错误码、两池全盘元数据、固定间隔重试时序和旧版同拓扑对照。旧版可能保留副本返回200,不要求它满足新增跨池删除契约。
有证据证明在目标操作前失败的 harness 尝试才能不计入有效运行,但仍计入总尝试。任何目标阶段的新503或数据不变量失败都不能通过补跑抹掉,必须暂停合并并定位。三次通过也只满足这项工程检查,不证明历史异常已消失;开放项在合并和正式发布评估时仍须可见。
## 8. 交付与恢复纪律
最终候选使用独立分支;main 的五个旧修改先保存完整内容、二进制补丁、SHA-256 和具名 Git 恢复引用,再核对原 HEAD/哈希及测试覆盖,只恢复这五个文件。禁止用整树 reset 或 clean 代替受控归一;若用户已有新增编辑,应保留并重新核对。
全量 cmd/internal、相关 race、构建、vet、lint、生成文件和兼容检查,以及上述 Linux 验收,绑定最终候选的实际 SHA。若纳入新的远端提交,重新记录基线与 head,复核变更并重跑受影响验收。文档与原始测试记录各自保留其对应版本,不篡改旧失败、不用新通过覆盖旧记录。
逐节点滚动升级未通过既有二进制校验检查,采用[协调停机方案](../bucket/lifecycle/access-tiering-removal.md#before-upgrading-a-build-with-access-tiering)。实验使用单个 Docker Linux VM 和 tmpfs;正式 tag、包、镜像、跨主机及生产部署仍是独立交付。未验证全部重复 UUID 或运行时统计守恒应如实披露,不能反向引入自动搬回/清理需求或无关重构。
## 9. 执行后记:最终基线、恢复窗口与验收
本次实际交付由 [PR #188](https://github.com/pgsty/silo/pull/188) 承载。选择的远端基线为 `89637554d60c27cfc51d2281d0a4fe15e415f06d`,移植没有包含本地独立 IAM/超时提交 `ebc9937d9`。前三项实现和历史文档逐提交通过 stable patch-id 对照;虚拟补回独立 IAM 差异后,完整树与原审查分支一致。
首次本地 lint 发现新增回调选择分支触发 `gocritic/ifElseChain`,因此追加等价的无表达式 `switch` 改写,保持 marker、指定版本和普通元数据查找的条件顺序及分支体。重新固定的代码候选为 [`41aa84609`](https://github.com/pgsty/silo/commit/41aa84609754769cfb1861d7fd060c2e84182b98)。这一提交上,全量 cmd/internal 得到 6,428 个测试及子测试通过、166 个跳过,50 个有测试的包通过;相关 race 得到 283 个测试及子测试通过。`make build`、全包构建、vet、lint、生成文件及 rebrand/compat 检查均通过。[Go CI](https://github.com/pgsty/silo/actions/runs/34925534139)、[DCO](https://github.com/pgsty/silo/actions/runs/34925534110)、[VulnCheck](https://github.com/pgsty/silo/actions/runs/34925534142) 和[发布流水线的测试运行](https://github.com/pgsty/silo/actions/runs/34925534198)共 11 项检查通过;后者没有发布正式制品。
第一次最终候选停池实验 `dcd5c2e5` 在源版本保全断言后失败:停掉另一池返回 `503 SlowDownRead`,源四盘版本保留;恢复后 DELETE 成功,节点 0/2 的 HEAD/GET 返回 404,节点 1 返回 503。DELETE 成功由脚本已通过的 204 断言确定,原输出没有单独保存该次 DELETE 响应。此次失败如实保留,不能把后来的成功写回原记录。
复核发现,`ListBuckets` 以及位于源池的现存版本 GET,不能证明每个协调节点对另一池的读取连接已经恢复。随后进行了旧基线与候选的同拓扑对照,探测的是从未写入过的随机 UUID,并且在这些探测之前没有执行任何 DELETE:
| 高时间分辨率对照 | 桶列表和现存版本 | 从未写入版本的 HEAD/GET | 观察到的恢复窗口 |
| --- | --- | --- | --- |
| `20502a2f`,旧基线 `89637554d` | 三节点均为 200 | 节点 0/2 为 404,节点 1 为 503 SlowDownRead,连续 16 组 | 从重启后的观察循环起算约 1.60–2.50 秒 |
| `246aaf77`,候选 `41aa84609` | 三节点均为 200 | 同样是节点 1 的 503,连续 11 组 | 约 1.67–2.30 秒 |
两边在缺失版本全节点连续三轮返回 404 后执行 DELETE,均得到 204、全节点 HEAD/GET 404、目标 UUID 在八盘均不存在且其他版本可读。另有两次较低时间分辨率诊断未捕捉到窗口,同样保留;这四次诊断不计入三次升级验收。
这给出了基线在零 DELETE 下的正向复现,证明原恢复条件不足。`getLatestObjectInfoWithIdx` 的读选择函数与基线文本相同:现存副本可以遮蔽另一池的读错误;缺失版本则必须确认所有池,不可读时返回 503。Claude 复核后同意修正实验准备条件并继续验收,明确反对把这一路径的 503 改成 404。最初失败没有瞬时 RPC 全貌,不能逐请求追溯每条连接;这些对照也不能给 `83676ca2` 归因。
修订后的验收在升级/恢复后逐节点探测从未写入的 UUID,记录首次全 404 时刻,要求连续三轮全 404;并列保存各节点 `admin info` 的全盘状态。最多等待 30 秒,超时仍失败,DELETE 之后仍严格要求 204/404,不把 503 纳入通过条件。后续失败保留实验资源供即时取证,再受控清理。
修订准备条件后的独立停池验收 `40a59b3b` 通过:离线 DELETE `503 SlowDownRead`,源四盘版本保留;恢复门约 1.48 秒完成,DELETE `204`,三节点 HEAD/GET `404`,目标在八盘均不存在,另一版本仍可读。恢复早期 `admin info` 中也记录到了各节点不同的离线盘视图,最后恢复为全盘正常。
完整升级验收随后实际进行了三次尝试:
| 尝试 | 运行 | 结果与证据边界 |
| --- | --- | --- |
| 1 | `83f88c59` | 准备失败,未启动候选。旧版 rebalance 报 Completed、搬移版本数为 0;源池占用约 6.5%,到平均空闲目标的差值约 3.13%,落入代码既有 5% 容差。增加造数从 64 到 192 个 2 MiB 版本后再试;本次仍计入五次总尝试上限。 |
| 2,第一轮有效运行 | `ea58d0c5` | 通过。实际 rebalance 中断产生跨八盘的重复版本;停机复制同一数据供旧版对照。旧版 DELETE204 后仍可读;候选 DELETE204 后四节点立即及后续固定间隔 HEAD/GET 均404,八盘目标清除、另一版本四节点可读;旧 ILM/生命周期编辑和真实缓存 v9→v8 通过。 |
| 3,第二轮有效运行 | `2059bc6b` | **候选失败,阻断合并。** 两阶段逐节点缺失版本连续三轮404、admin info全盘ok之后,DELETE明确204(约12.98ms);节点2随后的HEAD503/GET503 SlowDownRead,另外三节点404;八盘快照均无目标,首次重试及后续采样全404,其他版本四节点可读。首次失败保留,不因重试恢复改判。 |
第三次尝试发生后停止剩余验收,保留原容器、卷、元数据与响应时序进行诊断。不能把四节点顺序探测中的“节点2异常”直接解释为永久节点故障:它也可能与采样时间有关。随后在保留环境中,对三个额外重复版本做并发、不同顺序的“刚删 UUID / 从未写入 UUID”对照,候选稳定期均为204/404,未复现;旧版克隆数据的双次 DELETE 对照则遇到 `SlowDownWrite`,没有完成其全部断言。这些是诊断结果,不补入有效升级通过计数,也不证明第三次尝试已解释。
为观察逐盘返回,另在独立临时工作树编译仅增加日志的诊断二进制,**没有进入 PR**。它证明了第二层准备检查盲点:`getObjectFileInfo` 的四个响应信号中,可以只有两个实际 `file version not found`,其余两个是被跳过的盘所保留的 `errDiskOngoingReq`;`objectQuorumFromMeta` 的预期读 quorum 为2,因而仍可返回404。同期 `admin info` 汇集各服务器本地盘态为ok,不能证明请求节点到各盘的路径都可用。一次诊断 DELETE 的逐盘返回为 `[nil, nil, drive not found, drive not found]`,达不到写 quorum 3,故返回 `SlowDownWrite`。
Claude 撤回了此前“逐节点三轮404已是最强全池准备条件”的表述,同意这只能证明读 quorum,不能证明全盘可达或写 quorum。诊断给出了候选机制,但**第三次尝试失败瞬间没有逐请求逐盘日志,仍不足以确定其具体原因**;不能把后来稳定期的成功、诊断中的配额不足,或对错误名的解释当成该次故障的直接证据。
## 10. 首次执行的停止位置与恢复资料
首次执行停止时,代码为 `41aa84609`;后续提交只记录执行,不改变已测试的生产代码。当时有效升级运行是一通过、一失败,未满足三次有效运行全部通过的约定;总尝试3次,没有通过继续补跑消耗剩余次数来冲淡失败。`2059bc6b` 和先前的 `83676ca2` 均保持 **OPEN**。Claude 与 Codex 当时的裁定是 **NO-GO for merge**,没有把证据不足升级为“已修复”“既有问题”或“暂态”。后续收尾见第 11 节;实际合并状态以 [#188](https://github.com/pgsty/silo/pull/188) 为准。
主工作区原五文件的完整内容、二进制补丁和 SHA-256 已归档;另有可达 Git 引用 `refs/archive/access-tiering-main-five-files-20260915`,指向快照 `c7fbfc6ada0f0f2abcbe8c0a9681f07e12dafe52`。逐文件核验快照与原已审内容一致。首次停止时因验收未通过,没有执行五文件恢复,也没有移动或改写独立 IAM 提交 `ebc9937d9`。当时唯一待交付候选在 PR 分支,旧变体冻结等待验收裁定。
本机执行资料归档在 `~/.codex/outputs/silo-access-revert-assessment-20260915/final-execution/`:保存了每次尝试、旧/新基线对照、二进制 SHA-256、准备条件、源/目的全盘元数据、诊断补丁、独立评审意见,以及受控清理记录。原始失败记录不覆写;保存的诊断卷内容用于继续调查,不是生产数据或发布制品。
继续推进需要一次能区分机制的取证:在条件可控的升级实验中,于失败请求当时记录每池每盘的真实应答和错误类型,并同时读刚删版本与从未写入版本;明确区分盘面不同步、请求节点的不可用路径与其他原因。若四盘均真实应答仍返回503,应沿错误归约/元数据路径定位;若盘不可用,应查明连接或初始化状态,并验证准备条件。后续成功本身不能关闭本次失败,更不能通过把不可判定的503改成404来满足验收。正式发布、制品和生产部署继续作为独立交付。
## 11. 收尾复核:准备检查、可控机制与独立新轮次
后续收尾没有继续修改生产代码。旧基线 `89637554d` 与候选 `41aa84609` 使用各自的独立诊断构建,仅对测试桶记录逐盘应答;这些日志补丁没有进入 PR。第一轮完整诊断 `53ebe161` 通过但未复现503,仅作为一个样本保留。第二轮 `c951f762` 得到了不同的、可直接解释的失败:两个池的删除调用均记录 `quorum=3 errs=[nil,nil,nil,drive not found]`,DELETE204、所有读样本404,但八盘快照的盘3、7仍保留目标版本。它符合既有写 quorum 契约,并正向证明旧准备门会放行尚未完成挂盘的协调节点;它没有复现或解释 `2059bc6b` 那一次请求。
审查还纠正了两项推断:后续诊断进程在04:19的重连日志不能用于解释03:56的原失败;错误归约按具体错误值计数,不能把“最高同值计数不足”简单等同于“实际应答盘数不足”。原失败之后的全盘快照也不是失败瞬间的原子快照。这些界限继续保留。
准备检查改用服务器已有的 storage trace:从每个 S3 节点,对各自唯一、从未写入的对象执行 `GetObjectTagging`,该路径等待所有盘;按唯一对象名和后端节点、盘路径匹配真实 `storage.ReadVersion` 应答。四节点双池需要每轮32条实际缺失应答,连续三轮完成才满足准备条件;单纯404或 `admin info` 的本地盘状态不够。探针不创建对象,不改变服务器读写语义。trace 输出是多行 JSON 对象流,按流解析;缺失 trace 证据会使检查失败,不能据此假定对应盘健康。
为了比较相同的跨池写入负载,停止真实 rebalance 夹具后复制两份相同数据。旧基线使用 #178 已有的 `If-Match` 条件 DELETE,候选使用普通指定版本 DELETE,二者都处理同一目标的两个池副本。两臂分别完成96次“已删版本/从未写入版本”的 HEAD/GET 对照,均404,目标八盘均清除。准备检查分别耗时约0.60秒和17.08秒。这是一对匹配样本,没有观察到候选专属差异,不是吞吐比较,也不能排除所有可能的故障机制。
`f55967b7` 另用明确制造的重复副本夹具完成因果实验,而非冒充真实 rebalance:pool0 的四盘由一个节点承载并保留 namespace 锁;pool1 的四盘分布到另外四个节点。先停止 pool1 一盘,DELETE204 后直接确认只有该盘残留目标版本。重新接入这份副本,再停止 pool1 两块已删除该版本的盘,保留“一盘返回版本、一盘返回缺失”的读视图。三个被测协调节点的已删版本均返回 `503 SlowDownRead`,从未写入版本均返回404;同一失败请求的逐盘记录明确为 `[drive not found, drive not found, file version not found, nil]`,没有足够的同值应答达到读 quorum。整个实验没有丢失 pool0 的 namespace 锁,也没有将不确定状态改为404。
该因果实验的部分节点重启未在35秒内恢复全部40条访问路径,原实验因此仍记 FAIL;正向机制观察与这个恢复失败分开记录。随后协调重启全部五节点,独立恢复检查通过,五节点 HEAD/GET 均404。这不是滚动恢复已通过的声明,也不追溯原 `2059bc6b` 的逐盘状态。可控复现确定的是故障机制类别,原单次实例继续 **OPEN**。
Claude 复核上述证据后同意显式开启 **R-Upgrade-2**:原轮次的一通过、一失败和总尝试3次保持原样,不合并计数,不静默重置。新轮使用未经诊断修改的 `41aa84609` 二进制,要求三次有效运行全部通过、最多五次总尝试;任何新503或数据不变量失败仍须停止。每次 DELETE 前后均检查32条路径,响应后的即时 HEAD/GET 先于后置准备检查执行,避免等待掩盖短暂错误。
| 新轮次运行 | 完整运行耗时 | 升级后32路径准备耗时 | 验收结果 |
| --- | --- | --- | --- |
| `8eb016db` | 90.20秒 | 17.04秒 | PASS |
| `2a2b7b64` | 80.83秒 | 0.62秒 | PASS |
| `371e7b9f` | 96.28秒 | 0.60秒 | PASS |
三次均实际走到候选阶段:DELETE204,所有即时及后续 HEAD/GET 样本404,目标在八盘均不存在,其他版本从四节点读回;旧配置、普通生命周期规则编辑及真实缓存 v9→v8 均通过。删除前后各三轮32路径检查也全部通过。准备时间有明显波动,应验证访问路径,不能用固定等待秒数代替检查。这些结果满足修正准备条件后的有界验收;不把有限样本写成“历史503已消失”,不把不同阶段的诊断通过计入新轮次。
新轮仍绑定生产代码 `41aa84609`(Linux 二进制 SHA-256 `28e1339d630a22fa5a0e4659b6182224e81f6cd7cd4079856534390e40a697b1`),后续仅更新文档。合并前须核对源码等价性和最终 CI,按第8节的恢复纪律处理旧五文件,保留独立 IAM 提交。原 `2059bc6b`/`83676ca2`、部分重启恢复边界、单 VM/tmpfs、未验证全部不同重复 UUID 和整份运行时统计守恒继续可见;不为此新增自动搬回、清理服务或读错误降级。
本轮详细资料位于原归档的 `final-execution/closure-20260915/`,包括匹配对照、同请求逐盘日志、`qualification-summary.json`、独立 R-Upgrade-2 账本、诊断补丁和完整卷归档。原失败现场、后续对照及恢复后的卷分别标注时点。临时实验资源在归档验证后清理;这些资料与正式发布制品区分管理。
File diff suppressed because it is too large Load Diff
-190
View File
@@ -1,190 +0,0 @@
# SILO stack: Go 1.27 compatibility audit
> This is a dated investigation, with the source and runtime boundaries recorded
> below. It does not establish the current dependency pins or a later release.
> See [the current changelog](../../CHANGELOG.md) and
> [component matrix](https://silo.pgsty.com/compatibility/versions/).
2026-09-09. Scope: the maintained Server, silo-pkg, mcli, and Console. This extends
the [OIDC #154 investigation](issue-154.md) to other paths using the same TLS
configuration and to adjacent standard-library changes. It records local
validation performed before commit. No pushes, issue comments, releases, or
production changes were made during the audit.
## Findings and changes
| Component | Finding | Change |
| --- | --- | --- |
| Server, baseline `d1105bbb3d4a0afa33b3a4ac11b821235038ed0e` | Eight TLS configuration sites explicitly use the same ML-KEM-containing curve list, overriding `tlsmlkem=0` in Go 1.27. | Remove the eight assignments and obsolete `TLSCurveIDs` helper. Retain Go defaults in HTTP clients, replication, cloud client-certificate transport, both grid links, etcd, and the S3 listener. Add wire-level regression tests. |
| silo-pkg, baseline `a92c54d` | No built-in explicit PQ curve list. Web environment transport uses defaults; LDAP/OIDC accept caller configuration. | Add a real web-environment TLS regression test and document runtime-default selection. Keep the Go 1.26 library floor. |
| mcli, baseline `fcd5cad8` | S3/Admin transport and alias/TOFU dialer use default curves. No instance of the Server's override bug found. | Add real S3/Admin transport and alias-dialer handshake tests; update Go/TLS upgrade notes. |
| Console, baseline `c103d08ec` | IdP, SILO/STS, Prometheus, and webhook clients share a transport using defaults. The HTTPS listener explicitly uses only P-256. | Add real IdP/SILO-client handshake tests; update Go/TLS upgrade notes. Retain the existing listener policy. |
The Server's general external HTTP transport is also used by OpenID discovery
and JWKS, identity plugins, notification/lambda checks, audit/log webhooks, and
S3 cloud backends. Fixing only the two OIDC callers would leave those other paths
affected. The broader correction supersedes the earlier OIDC-only candidate.
The retained defaults follow Go's implementation rather than duplicating its
`GODEBUG` parser. Go 1.27 explicitly changed the interaction between manually
selected curves and the `tlsmlkem`/`tlssecpmlkem` default controls. See the
[Go TLS release notes](https://go.dev/doc/go1.27#crypto/tls).
With no opt-out, default curves additionally include SecP256r1MLKEM768 and
SecP384r1MLKEM1024. With `GODEBUG=tlsmlkem=0`, hybrid exchanges are disabled for
default-configured TLS throughout the process. This is an intentional change
from the old fixed subset. `GODEBUG=tlssecpmlkem=0` disables just the SecP
hybrids while retaining X25519MLKEM768. TLS versions, cipher-suite policy,
certificate and hostname checks, client certificates, proxies, and HTTP/2 choices are not
relaxed by this patch. No automatic fallback after a TLS error is introduced.
## Reproduction and regression evidence
Before changing product code, the new Server test failed for all five tested
outbound constructors with `tlsmlkem=0`: general external HTTP, internode HTTP,
replication, cloud client certificates, and etcd. The server observed
`[X25519MLKEM768 X25519 P256 P384 P521]` in every case. Both TLS 1.2 and TLS 1.3
peers reproduced the problem. The inbound Server also accepted a PQ-only client
despite the same opt-out.
After the change, those tests pass on darwin/arm64 and linux/arm64. They exercise
20 outbound combinations (five constructors, two peer TLS versions, two debug
settings), plus inbound classical/PQ-only peers with the opt-out enabled and
disabled. The inbound opt-out case rejects a PQ-only peer while continuing to
accept P-256. The etcd test exercises its actual TLS configuration, not an etcd
cluster; the client-certificate test exercises construction/loading, not a full
mutual-authentication service.
The mc, Console, and silo-pkg handshake tests pass without product code changes.
All use a trusted synthetic certificate and inspect a real ClientHello. They
check both disabling and retaining ML-KEM. Console also retains its existing
unknown-CA, hostname, and endpoint-scoping regression checks.
A freshly source-built complete Linux Server then passed three isolated
integration scenarios using the existing synthetic IdP fixture:
| Scenario | Result |
| --- | --- |
| ML-KEM-intolerant IdP + `tlsmlkem=0` | Discovery/JWKS, IAM, Console, and authenticated Admin calls succeed; login 204 and bucket list 200. |
| Normal TLS 1.3 IdP, no opt-out | Same successful login, token exchange, STS, session, and bucket-list chain. |
| Add OIDC to a running Server | Actual Admin API accepts the provider under the compatibility setting. |
Both login scenarios reject modified JWT signatures and a wrong audience: no
session cookie is issued and authenticated bucket access returns 403. Login
currently reports these authentication failures as 500; that existing error
mapping is outside this TLS change. Curl in the same namespace returns HTTP/2
200. The containers use a pre-existing generic Debian image, `--network none`,
and no published ports; no Server or Console image was downloaded or used.
The fixture drives real Console HTTP APIs, not rendered browser interaction.
This is a conditional interoperability reproduction, not proof of the actual
customer ingress behavior. Grid's existing tests pass, but a production-style
distributed TLS cluster, external cloud providers, and real etcd/LDAP/Keycloak
deployments were not exercised. Disabling ML-KEM does not disable new ML-DSA
signature offers and does not repair an ingress that rejects those offers.
## Other Go changes checked
### macOS root certificates: a confirmed upgrade-visible change
Using the same public synthetic CA and `certs.GetRootCAs`, fresh-process probes
produce the following results when `SSL_CERT_FILE` points at that CA and
`SSL_CERT_DIR` points at an empty directory:
| Compiler / main module | CA from environment trusted? | Explicit CA argument trusted? |
| --- | --- | --- |
| Go 1.26.5 / `go 1.26.0` | No | Yes |
| Go 1.27.1 / `go 1.26.0` | No | Yes |
| Go 1.27.1 / `go 1.27.1` | Yes | Yes |
The Go 1.26 module built by Go 1.27 carries
`DefaultGODEBUG=...x509sslcertoverrideplatform=0`; explicitly setting that option
to `1` enables the new behavior. A Go 1.27 application can explicitly set it to
`0` to recover the previous platform behavior. The diagnostic source is
[cert-roots.go](issue-154/cert-roots.go). The consuming application's defaults
apply to library calls as well, so silo-pkg's older `go` directive does not
prevent the behavior in Server, mc, or Console.
This is expected standard-library behavior, not a reason to silently discard
configured CA variables or skip verification. Setting either variable replaces
Keychain trust with on-disk roots and Go's verifier; stale or incomplete paths
can break previously trusted connections. Unset inherited variables to restore
Keychain trust; explicit additional CAs still work. The three application
READMEs and the package README now document this. The package's Windows loader enumerates
the native ROOT store directly and does not call `SystemCertPool`; it does not
inherit this particular new setting. Windows behavior was reviewed in source,
not runtime-tested.
### JSON, HTTP, timers, and other compatibility controls
- **JSON:** checked owned JSON error handling and exercised policy/condition,
config, and authentication tests. `quick` uses typed `SyntaxError` and
`UnmarshalTypeError`; no owned decision depending on changed standard JSON
error text was found. silo-pkg's full suite passes with both Go compilers;
mc's command and Console's API/auth suites pass under Go 1.27. No broad
`nojsonv2` opt-out or serialization rewrite is justified by these results.
- **HTTP response closing:** Go 1.27's standard-library drain is bounded at
256 KiB and 50 ms. mc's actual early-close/cancel regression passes, including
the compressed S3 Select stream. Existing owned drain helpers can still have
independent timeout concerns; those are not newly caused by this Go change.
- **ALPN and custom connections:** mc's TLS dialer returns a real `*tls.Conn`;
its deadline wrapper is on the TCP dial path. No new accidental HTTP/2 opt-in
from the expanded `ConnectionState` interface support was identified in
these transports. Console keeps its existing transport policy.
- **Removed switches:** no maintained runtime/config reliance on `tlsrsakex`,
`tls3des`, `tls10server`, `x509keypairleaf`, or `asynctimerchan` was found.
Existing explicit cipher policies continue to be explicit. Timer uses in
package certificate reload and license refresh do not depend on buffered
timer-channel length/capacity. Unix EOF error changes do not expose a matching
owned error-type assumption on the reviewed Unix-socket paths.
- **Platform floor:** application READMEs now identify macOS 13 as the minimum
for Go 1.27 binaries. No Windows, old macOS, or PowerPC runtime claim is made.
## Validation and delivery
- Server: focused TLS tests on macOS and Linux; macOS race run; crypto, HTTP,
grid, and all `internal/config/...` tests; full Linux Server build and the
three integration scenarios above.
- silo-pkg: `make test` (lint and full race suite) on Go 1.27.1; full suite on
Go 1.26.5, preserving the library floor.
- mc: complete `./cmd` suite, including the new TLS test and existing S3 Select
early-cancel test.
- Console: complete `./api/... ./pkg/...` suites, including identity and STS
validation and the new handshake tests.
Go lint reports zero issues in all four repositories. Server's optional spelling
check is skipped because `typos` is not installed. Its lint target used the
already-installed matching golangci-lint v2.13.1 after a redundant download was
stopped; mc's lint was rerun serially after the global linter lock prevented the
first concurrent attempt. These tooling retries did not require source changes.
An independent Fable 5.1 Max review reproduced the old-code failures and ran the
new TLS tests with the race detector. Its shuffled mc and Console suites passed.
One shuffled silo-pkg run failed the untouched `certs.TestValidPairAfterWrite`;
that test passed three plain reruns, the same shuffle seed, and the full certs
package rerun. The certs package runs in a separate test binary from the changed
env tests; this was classified as an existing timing flake.
The review also found that committing the diagnostic fixture made rebrand CI
count synthetic IdP paths as product routes. The guard now excludes
`docs/investigations/`; the product compatibility baseline is unchanged.
The full proposed file set passes the guard, and a negative control adding a
product route still fails. After the review corrections, a fresh Linux Server
build passed all three integration scenarios again; the new binary identity and
rerun results are recorded in the evidence file's `post_review_validation`.
All source and dependency choices remain those of the maintained PGSTY stack.
`go.mod` and `go.sum` are unchanged in every repository. The Server uses its
existing pinned Console/pkg/mc modules; the companion changes are tests and
documentation, so no replacement graph or unpublished dependency version is
needed to build the runtime fix.
Validation used separate worktrees for Server, silo-pkg, mc, and Console, leaving
the original repositories and their `main` branches untouched. Build identities,
fixture results, and command outcomes are recorded in
[go127-stack-evidence.json](go127-stack-evidence.json). Absolute paths and branch
names in the captured evidence identify the environment at recording time.
The recorded module graph identifies the dependencies selected for those
builds; `go.sum` also retains checksums for unselected versions. Local image IDs
and temporary paths are historical evidence, not portable setup instructions.
@@ -1,534 +0,0 @@
{
"date": "2026-09-11",
"runs": [
{
"runid": "silo-v1-0806-9c53f14c",
"version": "0806",
"image": "sha256:29a498b24669cae1fed11c1a2fb2b3d73c68829a0a9c0b14e71b386671d38fac",
"nodes": 4,
"drives": 4,
"filesystem": "Linux tmpfs named volumes, held mounted across server restarts",
"drive_bytes": 268435456,
"status": "PASS",
"phases": [
{
"phase": "startup-admin",
"seconds_from_start": 4.157,
"first_online_by_coordinator": [
4.037,
4.078,
4.118,
3.737
]
},
{
"phase": "startup-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.239,
"seconds_from_start": 4.396,
"first_put_seconds_after_admin_by_coordinator": [
0.096,
0.116,
0.127,
0.138
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.162,
0.194,
0.217,
0.239
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "full-restart-admin",
"seconds_from_start": 2.084,
"first_online_by_coordinator": [
1.937,
1.976,
2.016,
1.643
]
},
{
"phase": "full-restart-canary",
"attempts": 45,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 14.449,
"seconds_from_start": 16.534,
"first_put_seconds_after_admin_by_coordinator": [
14.315,
13.573,
13.953,
0.05
],
"first_four_reads_seconds_after_admin_by_coordinator": [
14.373,
14.396,
14.426,
14.449
],
"transient_error_count": 129,
"first_transient_errors": [
{
"attempt": 1,
"operation": "put",
"node": 0,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 1,
"operation": "put",
"node": 1,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 1,
"operation": "put",
"node": 2,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 2,
"operation": "put",
"node": 0,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
}
]
},
{
"phase": "readback-15s",
"seconds_after_canary": 17.596,
"objects": 55,
"reads": 220,
"errors": []
},
{
"phase": "readback-30s",
"seconds_after_canary": 31.357,
"objects": 55,
"reads": 220,
"errors": []
},
{
"phase": "readback-60s",
"seconds_after_canary": 61.566,
"objects": 55,
"reads": 220,
"errors": []
},
{
"phase": "one-node-outage",
"existing_read": true,
"put": true,
"readers": 3
},
{
"phase": "rejoin-admin",
"seconds_from_start": 1.382,
"first_online_by_coordinator": [
1.27,
1.307,
1.343,
1.382
]
},
{
"phase": "rejoin-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.158,
"seconds_from_start": 1.541,
"first_put_seconds_after_admin_by_coordinator": [
0.013,
0.029,
0.039,
0.05
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.077,
0.102,
0.132,
0.158
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "final-readback",
"seconds_after_canary": 1.71,
"objects": 60,
"reads": 240,
"errors": []
}
],
"acknowledged_objects": 60,
"version_ids_recorded": true,
"sha256_recorded": true
},
{
"runid": "silo-v1-0903-a457573c",
"version": "0903",
"image": "sha256:b616a0cf8cb281e7e6bb3c9b1fb53875b4016a2878223925541c18f82d6c5ca3",
"nodes": 4,
"drives": 4,
"filesystem": "Linux tmpfs named volumes, held mounted across server restarts",
"drive_bytes": 268435456,
"status": "PASS",
"phases": [
{
"phase": "startup-admin",
"seconds_from_start": 4.271,
"first_online_by_coordinator": [
4.158,
4.194,
3.827,
3.867
]
},
{
"phase": "startup-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.2,
"seconds_from_start": 4.471,
"first_put_seconds_after_admin_by_coordinator": [
0.054,
0.068,
0.081,
0.092
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.118,
0.146,
0.167,
0.2
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "full-restart-admin",
"seconds_from_start": 2.284,
"first_online_by_coordinator": [
2.143,
2.193,
0.849,
0.902
]
},
{
"phase": "full-restart-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.191,
"seconds_from_start": 2.476,
"first_put_seconds_after_admin_by_coordinator": [
0.016,
0.029,
0.044,
0.057
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.093,
0.124,
0.158,
0.192
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "readback-15s",
"seconds_after_canary": 15.235,
"objects": 8,
"reads": 32,
"errors": []
},
{
"phase": "readback-30s",
"seconds_after_canary": 30.356,
"objects": 8,
"reads": 32,
"errors": []
},
{
"phase": "readback-60s",
"seconds_after_canary": 60.192,
"objects": 8,
"reads": 32,
"errors": []
},
{
"phase": "one-node-outage",
"existing_read": true,
"put": true,
"readers": 3
},
{
"phase": "rejoin-admin",
"seconds_from_start": 0.903,
"first_online_by_coordinator": [
0.791,
0.829,
0.865,
0.903
]
},
{
"phase": "rejoin-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.141,
"seconds_from_start": 1.044,
"first_put_seconds_after_admin_by_coordinator": [
0.012,
0.022,
0.035,
0.045
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.069,
0.096,
0.12,
0.141
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "final-readback",
"seconds_after_canary": 0.301,
"objects": 13,
"reads": 52,
"errors": []
}
],
"acknowledged_objects": 13,
"version_ids_recorded": true,
"sha256_recorded": true
},
{
"runid": "silo-v1-current-6a14eb18",
"version": "current",
"image": "pgsty/d12a:build",
"image_id": "sha256:307af7711e2e04ab75759cb42a1eef45c43c4404894c0e30dd19f742b107b922",
"platform": "linux/arm64",
"nodes": 4,
"drives": 4,
"filesystem": "Linux tmpfs named volumes, held mounted across server restarts",
"drive_bytes": 268435456,
"status": "PASS",
"binary_sha256": "1e4cd7b78ecfa1b0cf28f1961d48a220a712c6be1fd3a01b60ee99aec62f04a4",
"source_revision": "b32f2d9dd01a383a9991d34ffe248063463a031d+pool-consistency",
"phases": [
{
"phase": "startup-admin",
"seconds_from_start": 7.374,
"first_online_by_coordinator": [
7.207,
6.677,
6.74,
6.8
]
},
{
"phase": "startup-canary",
"attempts": 1,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 0.479,
"seconds_from_start": 7.854,
"first_put_seconds_after_admin_by_coordinator": [
0.266,
0.279,
0.292,
0.305
],
"first_four_reads_seconds_after_admin_by_coordinator": [
0.336,
0.378,
0.431,
0.479
],
"transient_error_count": 0,
"first_transient_errors": []
},
{
"phase": "full-restart-admin",
"seconds_from_start": 2.461,
"first_online_by_coordinator": [
2.264,
1.414,
2.396,
2.461
]
},
{
"phase": "full-restart-canary",
"attempts": 36,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 14.489,
"seconds_from_start": 16.951,
"first_put_seconds_after_admin_by_coordinator": [
0.013,
0.025,
14.301,
0.046
],
"first_four_reads_seconds_after_admin_by_coordinator": [
14.358,
14.394,
14.435,
14.489
],
"transient_error_count": 104,
"first_transient_errors": [
{
"attempt": 1,
"operation": "put",
"node": 2,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 1,
"operation": "get",
"node": 2,
"key": "full-restart-attempt-1-node-0",
"error": "An error occurred (SlowDownRead) when calling the GetObject operation (reached max retries: 0): Resource requested is unreadable, please reduce your request rate"
},
{
"attempt": 1,
"operation": "get",
"node": 2,
"key": "full-restart-attempt-1-node-3",
"error": "An error occurred (SlowDownRead) when calling the GetObject operation (reached max retries: 0): Resource requested is unreadable, please reduce your request rate"
},
{
"attempt": 2,
"operation": "put",
"node": 2,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
}
]
},
{
"phase": "readback-15s",
"started_seconds_after_canary": 15.0,
"seconds_after_canary": 22.006,
"objects": 113,
"reads": 452,
"errors": []
},
{
"phase": "readback-30s",
"started_seconds_after_canary": 30.003,
"seconds_after_canary": 36.571,
"objects": 113,
"reads": 452,
"errors": []
},
{
"phase": "readback-60s",
"started_seconds_after_canary": 60.0,
"seconds_after_canary": 65.743,
"objects": 113,
"reads": 452,
"errors": []
},
{
"phase": "one-node-outage",
"existing_read": true,
"put": true,
"readers": 3
},
{
"phase": "rejoin-admin",
"seconds_from_start": 2.49,
"first_online_by_coordinator": [
1.165,
2.417,
1.646,
2.087
]
},
{
"phase": "rejoin-canary",
"attempts": 37,
"puts": 4,
"gets": 16,
"hard_deadline_seconds": 60,
"gate_seconds": 13.977,
"seconds_from_start": 16.468,
"first_put_seconds_after_admin_by_coordinator": [
0.013,
0.024,
0.035,
13.889
],
"first_four_reads_seconds_after_admin_by_coordinator": [
13.909,
13.927,
13.953,
13.977
],
"transient_error_count": 36,
"first_transient_errors": [
{
"attempt": 1,
"operation": "put",
"node": 3,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 2,
"operation": "put",
"node": 3,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 3,
"operation": "put",
"node": 3,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
},
{
"attempt": 4,
"operation": "put",
"node": 3,
"error": "An error occurred (SlowDownWrite) when calling the PutObject operation (reached max retries: 0): Resource requested is unwritable, please reduce your request rate"
}
]
},
{
"phase": "final-readback",
"started_seconds_after_canary": 0.0,
"seconds_after_canary": 6.507,
"objects": 226,
"reads": 904,
"errors": []
}
],
"acknowledged_objects": 226,
"version_ids_recorded": true,
"sha256_recorded": true
}
]
}
-327
View File
@@ -1,327 +0,0 @@
#!/usr/bin/env python3
"""Bounded four-container restart/readback acceptance for pgsty/silo#116."""
import argparse
import hashlib
import json
import os
from pathlib import Path
import secrets
import shutil
import signal
import socket
import subprocess
import time
import uuid
from concurrent.futures import ThreadPoolExecutor
import boto3
from botocore.config import Config
ROOT = None
MCLI = None
CURRENT_BINARY = None
SOURCE_REVISION = None
IMAGES = {
'0806': 'pgsty/silo:RELEASE.2026-08-06T00-00-00Z',
'0903': 'pgsty/silo:RELEASE.2026-09-03T13-18-01Z',
'current': 'pgsty/d12a:build',
}
def docker(*args, timeout=40, check=True):
return subprocess.run(['docker', *args], check=check, capture_output=True, text=True, timeout=timeout)
class Deadline(BaseException):
pass
def bounded(seconds, operation):
def expired(*_):
raise Deadline(f'hard deadline of {seconds}s exceeded')
old = signal.signal(signal.SIGALRM, expired)
signal.setitimer(signal.ITIMER_REAL, seconds)
try:
return operation(time.monotonic() + seconds)
finally:
signal.setitimer(signal.ITIMER_REAL, 0)
signal.signal(signal.SIGALRM, old)
def remaining(deadline):
value = deadline - time.monotonic()
if value <= 0:
raise Deadline('absolute deadline exceeded')
return value
def run(version):
image_info = json.loads(docker('image', 'inspect', IMAGES[version]).stdout)[0]
runid = f'silo-v1-{version}-{uuid.uuid4().hex[:8]}'
out = ROOT / runid
out.mkdir(mode=0o700)
user, password = 'local116', secrets.token_urlsafe(24)
envfile = out / 'credentials.env'
envfile.write_text(f'MINIO_ROOT_USER={user}\nMINIO_ROOT_PASSWORD={password}\nMINIO_CI_CD=1\nMINIO_BROWSER=off\nGOMAXPROCS=2\n')
envfile.chmod(0o600)
env = {k: v for k, v in os.environ.items() if not k.startswith(('MINIO_', 'SILO_', 'MC_'))
and k.lower() not in {'http_proxy', 'https_proxy', 'all_proxy', 'no_proxy'}}
nodes = [f'{runid}-n{i}' for i in range(4)]
sockets = [socket.socket() for _ in nodes]
for sock in sockets:
sock.bind(('127.0.0.1', 0))
ports = [sock.getsockname()[1] for sock in sockets]
for sock in sockets:
sock.close()
volumes = [n + '-data' for n in nodes]
endpoints, ledger = [], []
result = {'runid': runid, 'version': version, 'image': IMAGES[version],
'image_id': image_info['Id'], 'platform': image_info['Os']+'/'+image_info['Architecture'],
'nodes': 4,
'drives': 4, 'filesystem': 'Linux tmpfs named volumes, held mounted across server restarts',
'drive_bytes': 268435456, 'phases': [], 'status': 'RUNNING'}
if version == 'current':
result['binary_sha256'] = hashlib.sha256(CURRENT_BINARY.read_bytes()).hexdigest()
result['source_revision'] = SOURCE_REVISION
bucket = 'canary-' + uuid.uuid4().hex[:10]
def save(event=None):
if event is not None:
result['phases'].append(event)
compact = {k: (len(v) if k in ('transient_errors', 'errors') else v) for k, v in event.items()}
print(json.dumps({'version': version, **compact}), flush=True)
(out / 'result.json').write_text(json.dumps(result, indent=2) + '\n')
(out / 'acknowledged.json').write_text(json.dumps(ledger, indent=2) + '\n')
def parallel(fn, values):
with ThreadPoolExecutor(max_workers=4) as pool:
return list(pool.map(fn, values))
def client(i, deadline):
timeout = remaining(deadline)
return boto3.client('s3', endpoint_url=endpoints[i], aws_access_key_id=user,
aws_secret_access_key=password, region_name='us-east-1',
config=Config(proxies={}, signature_version='s3v4', s3={'addressing_style': 'path'},
retries={'total_max_attempts': 1}, connect_timeout=timeout,
read_timeout=timeout, request_checksum_calculation='when_required',
response_checksum_validation='when_required'))
def admin_gate(origin, phase):
def check(deadline):
first = [None] * 4
while True:
states = []
for i in range(4):
try:
p = subprocess.run([str(MCLI), '--config-dir', str(out / 'mcli'), '--json',
'admin', 'info', f'n{i}'], env=env, capture_output=True,
text=True, timeout=min(5, remaining(deadline)))
info = json.loads(p.stdout)['info']
servers = info['servers']
ok = len(servers) == 4 and all(s.get('state') == 'online' and s.get('drives')
and all(d.get('state') == 'ok' for d in s['drives'])
and all(v == 'online' for v in s.get('network', {}).values()) for s in servers)
if ok:
(out / f'{phase}-admin-{i}.json').write_text(json.dumps(info, indent=2) + '\n')
if first[i] is None:
first[i] = round(time.monotonic() - origin, 3)
states.append(ok)
except (Exception,):
states.append(False)
if all(states):
event = {'phase': phase + '-admin', 'seconds_from_start': round(time.monotonic()-origin, 3),
'first_online_by_coordinator': first}
save(event)
return time.monotonic()
time.sleep(min(.25, remaining(deadline)))
return bounded(90, check)
def canary(phase, origin, admin_time, setup=False):
def check(deadline):
started = time.monotonic()
first_put, first_reads = [None] * 4, [None] * 4
errors, attempt, setup_done = [], 0, not setup
result['active_canary'] = {'phase': phase, 'errors': errors}
while True:
attempt += 1
remaining(deadline)
if not setup_done:
try:
try:
client(0, deadline).create_bucket(Bucket=bucket)
except Exception as e:
if 'BucketAlreadyOwnedByYou' not in str(e):
raise
client(0, deadline).put_bucket_versioning(Bucket=bucket, VersioningConfiguration={'Status': 'Enabled'})
setup_done = True
except Exception as e:
errors.append({'attempt': attempt, 'operation': 'setup', 'error': str(e)[:250]})
time.sleep(min(.25, remaining(deadline)))
continue
acked = []
for i in range(4):
key = f'{phase}-attempt-{attempt}-node-{i}'
payload = (key + '\n').encode() * 16384
try:
vid = client(i, deadline).put_object(Bucket=bucket, Key=key, Body=payload)['VersionId']
entry = {'phase': phase, 'key': key, 'version': vid, 'bytes': len(payload),
'sha256': hashlib.sha256(payload).hexdigest(), 'writer': i,
'ack_seconds_from_start': round(time.monotonic()-origin, 3)}
ledger.append(entry)
save() # Persist every acknowledged write, including failed rounds.
acked.append(entry)
if first_put[i] is None:
first_put[i] = round(time.monotonic()-admin_time, 3)
except Exception as e:
errors.append({'attempt': attempt, 'operation': 'put', 'node': i, 'error': str(e)[:250]})
reads = 0
for i in range(4):
own_reads = 0
for entry in acked:
try:
verify(i, entry, deadline)
reads += 1
own_reads += 1
except Exception as e:
errors.append({'attempt': attempt, 'operation': 'get', 'node': i,
'key': entry['key'], 'error': str(e)[:250]})
if own_reads == 4 and first_reads[i] is None:
first_reads[i] = round(time.monotonic()-admin_time, 3)
remaining(deadline)
if len(acked) == 4 and reads == 16:
result.pop('active_canary', None)
save({'phase': phase + '-canary', 'attempts': attempt, 'puts': 4, 'gets': 16,
'hard_deadline_seconds': 60, 'gate_seconds': round(time.monotonic()-started, 3),
'seconds_from_start': round(time.monotonic()-origin, 3),
'first_put_seconds_after_admin_by_coordinator': first_put,
'first_four_reads_seconds_after_admin_by_coordinator': first_reads, 'transient_errors': errors})
return time.monotonic()
(out / f'{phase}-canary-errors.json').write_text(json.dumps(errors, indent=2) + '\n')
if attempt == 1:
print(json.dumps({'version': version, 'phase': phase, 'first_attempt_errors': errors}), flush=True)
time.sleep(min(.25, remaining(deadline)))
return bounded(60, check)
def verify(i, entry, deadline):
got = client(i, deadline).get_object(Bucket=bucket, Key=entry['key'], VersionId=entry['version'])
try:
data = got['Body'].read()
finally:
got['Body'].close()
assert len(data) == entry['bytes'] and hashlib.sha256(data).hexdigest() == entry['sha256'], entry['key']
assert got.get('VersionId') == entry['version'], entry['key']
remaining(deadline)
def readback(label, origin):
def check(deadline):
started = time.monotonic()
errors = []
for entry in ledger:
for i in range(4):
try:
verify(i, entry, deadline)
except Exception as e:
errors.append({'key': entry['key'], 'node': i, 'error': str(e)[:250]})
save({'phase': label, 'started_seconds_after_canary': round(started-origin, 3),
'seconds_after_canary': round(time.monotonic()-origin, 3),
'objects': len(ledger), 'reads': len(ledger)*4, 'errors': errors})
return errors
return bounded(60, check)
try:
docker('network', 'create', runid)
for volume in volumes:
docker('volume', 'create', '--driver', 'local', '--opt', 'type=tmpfs',
'--opt', 'device=tmpfs', '--opt', 'o=size=256m', volume)
mounts = [arg for i, v in enumerate(volumes) for arg in ('--mount', f'type=volume,source={v},target=/keep/{i}')]
docker('run', '-d', '--pull=never', '--network', 'none', '--name', runid + '-keeper',
*mounts, '--entrypoint', 'sleep', 'alpine:3.23', '1800')
urls = [f'http://{n}:9000/data' for n in nodes]
def create(i):
extra = []
if version == 'current':
extra = ['--mount', f'type=bind,source={CURRENT_BINARY},target=/lab/silo,readonly',
'--entrypoint', '/lab/silo']
docker('create', '--pull=never', '--name', nodes[i], '--network', runid,
'--hostname', nodes[i], '--cpus', '2', '--memory', '3g', '--env-file', str(envfile),
'--mount', f'type=volume,source={volumes[i]},target=/data',
'-p', f'127.0.0.1:{ports[i]}:9000', *extra, image_info['Id'], 'server', '--address', ':9000',
'--console-address', ':9001', *urls)
parallel(create, range(4))
start = time.monotonic()
parallel(lambda n: docker('start', n), nodes)
for i, n in enumerate(nodes):
endpoint = 'http://' + docker('port', n, '9000/tcp').stdout.strip()
endpoints.append(endpoint)
env[f'MC_HOST_n{i}'] = endpoint.replace('http://', f'http://{user}:{password}@')
admin = admin_gate(start, 'startup')
canary('startup', start, admin, setup=True)
parallel(lambda n: docker('stop', '-t', '10', n), nodes)
start = time.monotonic()
parallel(lambda n: docker('start', n), nodes)
admin = admin_gate(start, 'full-restart')
gate = canary('full-restart', start, admin)
read_errors = []
for after in (15, 30, 60):
time.sleep(max(0, gate + after - time.monotonic()))
read_errors.extend(readback(f'readback-{after}s', gate))
assert not read_errors, 'acknowledged object readback failure; see result.json'
docker('stop', '-t', '10', nodes[3])
def outage(deadline):
verify(0, ledger[0], deadline)
payload = b'acknowledged with one Linux node offline' * 16384
key = 'one-node-outage'
vid = client(0, deadline).put_object(Bucket=bucket, Key=key, Body=payload)['VersionId']
entry = {'phase': 'outage', 'key': key, 'version': vid, 'bytes': len(payload),
'sha256': hashlib.sha256(payload).hexdigest(), 'writer': 0}
ledger.append(entry)
save()
for i in range(3):
verify(i, entry, deadline)
save({'phase': 'one-node-outage', 'existing_read': True, 'put': True, 'readers': 3})
bounded(60, outage)
start = time.monotonic()
docker('start', nodes[3])
admin = admin_gate(start, 'rejoin')
gate = canary('rejoin', start, admin)
assert not readback('final-readback', gate)
result['status'] = 'PASS'
except BaseException as e:
result['status'] = 'FAIL'
result['error'] = f'{type(e).__name__}: {e}'
raise
finally:
save()
for n in nodes:
log = docker('logs', n, check=False)
(out / (n + '.log')).write_text(log.stdout + log.stderr)
docker('rm', '-f', n, check=False)
docker('rm', '-f', runid + '-keeper', check=False)
for v in volumes:
docker('volume', 'rm', v, check=False)
docker('network', 'rm', runid, check=False)
envfile.unlink(missing_ok=True)
print(json.dumps({'version': version, 'status': result['status'], 'evidence': str(out)}), flush=True)
if __name__ == '__main__':
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('versions', nargs='+', choices=list(IMAGES))
parser.add_argument('--output', type=Path, required=True, help='directory for retained evidence')
parser.add_argument('--mcli', default=shutil.which('mcli'), help='native mcli executable')
parser.add_argument('--current-binary', type=Path, help='Linux binary matching the Docker architecture')
parser.add_argument('--source-revision', help='Git revision of --current-binary')
parser.add_argument('--current-image', default=IMAGES['current'], help='cached Linux base image for current binary')
args = parser.parse_args()
if not args.mcli or not Path(args.mcli).is_file():
parser.error('--mcli must point to an executable file')
if 'current' in args.versions and (not args.current_binary or not args.current_binary.is_file()):
parser.error('current requires --current-binary')
ROOT = args.output.resolve()
ROOT.mkdir(mode=0o700, parents=True, exist_ok=True)
MCLI = Path(args.mcli).resolve()
CURRENT_BINARY = args.current_binary.resolve() if args.current_binary else None
SOURCE_REVISION = args.source_revision
IMAGES['current'] = args.current_image
for version in args.versions:
run(version)
-409
View File
@@ -1,409 +0,0 @@
# SILO #154:OIDC discovery 连接重置调查
> This is a dated investigation, with the source and runtime boundaries recorded
> below. It does not establish the current dependency pins or a later release.
> See [the current changelog](../../CHANGELOG.md) and
> [component matrix](https://silo.pgsty.com/compatibility/versions/).
前两轮调查时间:2026-09-09;公开 issue 最后核对于 07:53 UTC。第二轮补充同源码、同依赖、不同 Go 工具链的 Linux 完整 Server 对照和候选补丁认证链路验证。
**后续更新:用户已授权扩展至整个 SILO 技术栈并修复。现已确认 Server 其他 TLS 路径也存在同类覆盖问题,并在产品工作区完成统一使用 Go 默认曲线的修复。当前实现、验证和交付状态见 [全栈调查](go127-stack.md)。下文保留前两轮的诊断与当时的 OIDC 局部候选;“未修改产品”和“不要扩大范围”等表述仅适用于当时的调查阶段,局部候选已被后续全路径修复取代。**
## 判断与处理顺序
**目前可以确认是 Server 发出的 discovery GET 失败,随后 IAM 初始化等待,Console 初始化也被延后;还不能确认真实连接由谁、在哪个协议阶段重置。优先调查 SILO/Go 客户端与 IdP 前置 TLS 终止器、代理或 WAF 的互操作。** 现有证据不支持将其定性为证书错误、Keycloak 配置错误、JWT 校验错误或 `coreos/go-oidc` 回归。
有两项与版本相关、可在本地验证的 TLS 差异:
1. **Go 1.27 改变了 `GODEBUG=tlsmlkem=0` 与显式 `CurvePreferences` 的关系。** SILO 两个版本都显式包含 X25519MLKEM768。旧版 Go 1.26.5 会根据该环境选项移除它;Go 1.27.1 保留显式配置。在模拟拒绝 ML-KEM 的入口上,能重现“相同选项下旧 Server 启动成功,新 Server 持续 reset”。**客户是否设置过这个选项尚未知,不能把条件性复现当成客户根因。**
2. **Go 1.27 的 ClientHello 新增 ML-DSA 签名算法。** 没有上述环境选项时,新旧版本也会发送不同的握手。人为拒绝新算法编号的入口同样能产生旧成功、新失败。真实入口是否存在这种行为尚未知。
另外,HTTP User-Agent 从 `MinIO` 变成了 `Silo`;若连接在 TLS 完成、GET 发出之后被重置,应优先查 WAF、User-Agent 规则和 HTTP 路由,而不是继续调整 TLS。
最短路径:**先在故障进程所在环境拿到实际 Go 版本、是否设置 `tlsmlkem=0`、目标 IP 和 TLS 完成与否;再针对已证实的分支处理。** 优先修正入口兼容性或错误路由。若确认是上述环境选项失效,只为 OpenID 出站请求恢复 Go 默认曲线选择,是目前最小的代码候选。现阶段不宜做全局 TLS 改动或整体依赖回退。
## 第二轮结论:已隔离 Go 因素,局部候选修复通过 Linux 验证
**可以复现一种由 Go 1.26.5 → 1.27.1 单独触发的兼容性回归。建议保留 Go、针对已确认分支修复;整体回退只用于临时恢复服务。** 这里的“已确认”指本地实验机制,仍不等于已经确认客户入口的根因。
四个完整 Server 都从本地源码编译为 `linux/arm64`,在现成通用 Debian 12 基础镜像的 `--network none` 容器中运行;Server、合成 IdP、Console 和 curl 共用同一 loopback 网络,没有暴露端口。没有下载或运行 Server/Console 镜像。旧源码为 `d88f46ccee345a9c2fabe2d221d9a9e56bc11aec`,当前源码为 `d1105bbb3d4a0afa33b3a4ac11b821235038ed0e`。
旧源码两次构建使用完全相同的 `go.mod`、`go.sum` 和实际链接模块版本,仅替换编译器;当前源码与候选补丁构建的依赖图也完全相同。实际二进制 SHA-256、`go version -m` 依赖和构建选项保存在 [Linux 证据](issue-154/linux-evidence.json)。
下表的入口**人为设置为见到 X25519MLKEM768 就发 TCP RST**,Server 都设置 `GODEBUG=tlsmlkem=0`:
| 源码与编译器 | Server 初始 ClientHello | 完整 Server 结果 | 同容器 curl |
| --- | --- | --- | --- |
| 同一份旧源码 + Go 1.26.5 | 275 字节,无 ML-KEM | discovery、JWKS、IAM、Console 正常;cluster 200 | HTTP/2 200 |
| 同一份旧源码 + Go 1.27.1 | 1509 字节,仍包含 ML-KEM | `connection reset by peer`;IAM 等待;cluster 503;Console 未启动 | HTTP/2 200 |
| 当前源码 + Go 1.27.1 | 1509 字节,仍包含 ML-KEM | 同样失败 | HTTP/2 200 |
| 当前源码 + 局部候选补丁 + Go 1.27.1 | 287 字节,无 ML-KEM | 完整启动及合成 OIDC 登录成功 | HTTP/2 200 |
这将该条件下的回归定位到工具链行为,而非 Console、pkg、mc 或 `coreos/go-oidc` 升级。Go 1.27 的发布说明明确将此作为有意改变:`tlsmlkem` / `tlssecpmlkem` 只控制默认曲线集合,显式指定的集合可以继续启用这些算法。SILO 现有曲线列表显式包含该算法,因而原来的兼容开关在这条路径失效。[Go 1.27 crypto/tls 说明](https://go.dev/doc/go1.27#crypto/tls)。
共完成 **12 个 Linux 场景**,其中失败用例按预期失败:
- 正常 TLS 1.2 / P-256 / RSA / AES-256-GCM 的 IdP,旧源码用两个 Go 版本编译均正常,证明不是 Go 1.27 普遍无法连接这套 TLS。
- 上表四个对照均符合预期;候选补丁如果不设置 `tlsmlkem=0`,仍被 ML-KEM 拒绝规则拦截。补丁恢复显式兼容选项的作用,不会自行关闭后量子算法。
- 候选补丁在上述 TLS 1.2 兼容场景、以及无 GODEBUG 的正常 TLS 1.3 场景,都完成 Console 登录信息获取 → IdP authorization redirect → callback → token exchange → STS 凭据 → Console 会话 → 桶列表读取。登录 API 为 204,桶列表为 200。
- 两个登录场景分别提交错误签名与错误 audience 的 JWT,登录返回 500、桶列表返回 403,没有生成会话 cookie。这里只确认认证未放行,没有将现有 500 状态码行为改为另一项修复。
- 不信任 CA 时,候选 Server 仍因 `x509: certificate signed by unknown authority` 停在 IAM 初始化。没有关闭证书校验。
- 不配置 OIDC 启动后,经实际 Admin API 添加同一合成 provider:当前源码失败且返回 reset;候选补丁成功。
- 入口改为拒绝 ML-DSA 签名编号,候选补丁加 `tlsmlkem=0` 仍失败。这是另一条机制,当前补丁没有解决它。
认证流程由 Python 驱动真实 Console HTTP API;IdP 使用一次性合成授权码和自行签发的实验 JWT,没有客户账户。没有执行浏览器页面交互、登出、真实 Keycloak 或生产数据升级测试。TLS 1.3 对照实际协商 TLS 1.3/P-256,不是所有新增后量子曲线的互操作覆盖。
### 修复与回退的取舍
[候选补丁](issue-154/openid-default-curves.patch) 只有三个文件:增加 OpenID 专用 transport helper,将其 `TLSClientConfig.CurvePreferences` 设为 `nil`,再替换 IAM 初始化和 OpenID 配置校验两个调用点。测试时仅应用于隔离源码副本,当前产品工作区未应用;Go 版本与所有依赖保持不变。它让 Go 默认策略和现有兼容开关接管外部 IdP 的密钥交换,其他 TLS 参数继续来自现有构造函数。
如果客户证据确认此分支,建议采用该局部修复,并在客户实际入口复验。如果客户没有设置 `tlsmlkem=0`,它不能单独解释旧版成功:旧版默认也发送 ML-KEM。此时应继续区分 ML-DSA、其他握手变化、HTTP/WAF 规则与网络路径;不把此补丁直接宣称为 #154 的完整修复。
**不建议将当前产品直接改回 Go 1.26。** 实测以 Go 1.26.7 和 `GOTOOLCHAIN=local` 读取当前源码即被 `go.mod requires go >= 1.27.1` 拒绝;所固定的 Server、Console、mc 均声明 Go 1.27.1。回退需要进一步调整这些模块及可能的传递依赖,不是只换一个编译器版本。这个报错证明当前依赖图不能原样回编,并不证明经过额外适配后绝对无法回编。报告中已恢复服务的旧版可作为临时运行状态,不能把这种处置等同于完成兼容修复。
第二轮的运行脚本是 [run-linux.py](issue-154/run-linux.py),结果是 [linux-evidence.json](issue-154/linux-evidence.json),具体构建和运行步骤见文末。没有 issue 评论、提交、推送、合并或发布。
候选副本另通过 Go 1.27.1、darwin/arm64、CGO 关闭的 `go test -mod=readonly -count=1 ./internal/http ./internal/config/identity/openid`;补丁可以干净应用到调查基线。这些包级测试与上述 Linux 集成验证分别记录,不将其当成 Linux 单元测试结果。
## 范围、版本与公开证据
- 初始工作区干净,处于 detached HEAD;调查分支为 `codex/investigate-oidc-154`,基线与远端 main 均为 `d1105bbb3d4a0afa33b3a4ac11b821235038ed0e`。
- 已读取 `/Users/vonng/pgsty/silo/AGENTS.md` 及工作区适用说明。维护范围为 PGSTY 的 Server、Console、mc、silo-pkg;上游 MinIO 仅作参考。
- 第一轮 Server 与 transport 探针为本地源码构建的 darwin/arm64 程序,第二轮完整 Server 交叉构建为 linux/arm64;均使用 `CGO_ENABLED=0 GOWORK=off`。fixture 与 Admin 辅助程序也由本机 Go 构建。历史源码使用 `git archive` 导入隔离临时目录。没有下载或运行 Server/Console Docker 镜像,没有访问客户端点,没有改动真实服务或数据,没有评论 issue、推送、合并或发布。
- 产品源码、`go.mod`、`go.sum` 未修改。本目录中的 Go 文件是显式运行的调查工具,带 `//go:build ignore`,不进入正常构建。
| 项目 | 旧版:2026-08-04 | 报告故障版:2026-09-03 | 调查时 main |
| --- | --- | --- | --- |
| 完整 tag | `RELEASE.2026-08-04T00-00-00Z` | `RELEASE.2026-09-03T13-18-01Z` | 无新 release 声明 |
| 源码 commit | `d88f46ccee345a9c2fabe2d221d9a9e56bc11aec` | `9b11dc9469e650815b775cb47b039610644f5da4` | `d1105bbb3d4a0afa33b3a4ac11b821235038ed0e` |
| go.mod / Docker 构建定义 | Go 1.26.5 | Go 1.27.1 | Go 1.27.1 |
| 本地 Server / 探针实际编译器 | Go 1.26.5 | Go 1.27.1 | Go 1.27.1 |
| Console replacement | `v0.0.0-20260804042150-b952a1202869` | `v0.0.0-20260903111932-464a59d73ada` | `v0.0.0-20260908142700-c103d08ec36a` |
| mc replacement | `v0.0.0-20260801042411-ad10a2a10b76` | `v0.0.0-20260903063637-a2ef95c035d9` | `v0.0.0-20260909015522-fcd5cad8247f` |
| PGSTY silo-pkg | v3.11.0,替换历史 minio/pkg 路径 | v3.13.2,直接依赖 | v3.13.3,直接依赖 |
| coreos/go-oidc/v3 | v3.17.0 | v3.21.0 | v3.21.0 |
| x/crypto | v0.54.0 | v0.56.0 | v0.56.0 |
| x/net | v0.57.0 | v0.58.0 | v0.58.0 |
| x/oauth2 | v0.36.0 | v0.36.0 | v0.36.0 |
版本依据:[旧版 go.mod](https://github.com/pgsty/silo/blob/d88f46ccee345a9c2fabe2d221d9a9e56bc11aec/go.mod)、[故障版 go.mod](https://github.com/pgsty/silo/blob/9b11dc9469e650815b775cb47b039610644f5da4/go.mod)、[本次 main go.mod](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/go.mod)。还核对了各 tag 的 `Dockerfile.goreleaser` 和 release workflow。它们声明相应 Go 构建版本、禁用 CGO;历史归档、本次探针和 Server 均用 `go version -m` 核实实际编译器与 replacement。**这些不是客户实际镜像 digest 或其内二进制的取证,后者仍需 `--version` / build info 确认。**
[Issue #154](https://github.com/pgsty/silo/issues/154) 当前 OPEN,最后更新时间 `2026-09-08T05:30:42Z`,评论数为 0。报告包含升级后 Server 初始化失败、旧版回退恢复、测试实例添加 OIDC 失败,以及 curl 成功的输出。
curl 那次连接验证了所收到的证书链,并选择 TLS 1.2、`ECDHE-RSA-AES256-GCM-SHA384`、P-256 和 HTTP/2;解析得到两个 IPv4,记录中实际访问了其中一个。**公开命令使用 `docker run --rm` 新建容器,并非 `docker exec` 进入原故障容器;同镜像不能证明同网络命名空间、环境变量、CA 挂载、DNS 结果或出口。** 域名、realm、IP 已脱敏,本次不推测其真实值。
## 实际请求链
### IAM 与配置校验
```text
Server startup
IAMSys.Init
openid.LookupConfig
parseDiscoveryDoc: GET .well-known/openid-configuration
PopulatePublicKey: GET discovery 中的 jwks_uri
IAM store 初始化
Console 初始化
Console 添加 OIDC
AdminClient.AddOrUpdateIDPConfig
Server addOrUpdateIDPHandler
validateConfig(identity_openid)
同一个 openid.LookupConfig / NewHTTPTransport
```
`parseDiscoveryDoc` 是 SILO 自己的实现,用标准库 `http.Client` 发 GET,收到成功响应后才解码 JSON;这次日志中的 `Get ... read tcp ... reset` 发生在该调用返回响应之前,不能进一步区分 TLS 与 HTTP。请求本身不需要客户 client secret、token 或私钥。[IAM 调用点](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/iam.go#L278-L288)、[discovery 实现](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/internal/config/identity/openid/jwt.go#L261-L285)、[JWKS 实现](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/internal/config/identity/openid/jwt.go#L89-L110)。
Console 的添加表单通过 Admin API 触发 Server 配置校验。校验成功才写入配置;本地已重现该 API 返回同样的 reset,恢复 IdP 后再次创建成功并要求重启。[Console 调用](https://github.com/pgsty/silo-console/blob/c103d08ec36a/api/admin_idp.go#L79-L114)、[Admin 客户端](https://github.com/pgsty/silo-console/blob/c103d08ec36a/api/client-admin.go#L593-L595)、[Server 校验和保存边界](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/admin-handlers-idp-config.go#L127-L150)、[OpenID 配置校验](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/config-current.go#L353-L359)。
`coreos/go-oidc.NewProvider` 在 Server 中的使用是 `MockOpenIDTestUserInteraction` 测试辅助函数,不是这次 IAM discovery 调用链。升级此依赖不能单独解释或修复该 GET。`crypto/tls` 和这里的 `net/http` 属于 Go 标准库,不能用 go.mod 中 `x/crypto`、`x/net` 的版本代替它们的实际行为。[辅助函数](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/utils.go#L994-L1012)。
### transport 参数与环境
| 项目 | 实际行为及意义 |
| --- | --- |
| 构造 | `NewHTTPTransport()` → `NewHTTPTransportWithTimeout(time.Minute)` → `xhttp.ConnSettings`。每次 IAM 初始化重试都会重新构造。 |
| TLS 版本 | 未显式设置 Min/Max;所比较 Go 工具链的正常默认范围是 TLS 1.2–1.3。证书、主机名验证开启。 |
| 密码套件 | 显式 `TLSCiphersBackwardCompatible()`,包含 curl 成功使用的 ECDHE-RSA/AES-256-GCM。没有理由为本 issue 添加旧 RSA、3DES 或 SHA-1 例外。 |
| 曲线 | 显式 `{X25519MLKEM768, P256, X25519, P384, P521}`,两个 release 和当前 main 相同。实际上线顺序还受 Go 实现控制,见实验。 |
| ALPN / HTTP | `EnableHTTP2=false`,同时设置自定义 TLS config 和 DialContext;实测 ClientHello 没有 ALPN,GET 使用 HTTP/1.1。三个版本一致。`GODEBUG=http2client=0` 对此基线路径无修复价值。 |
| 代理 | `http.ProxyFromEnvironment`;HTTPS URL 使用 `HTTPS_PROXY` / `https_proxy` 与 `NO_PROXY` / `no_proxy`。这两版标准库均优先非空大写值,不使用 `ALL_PROXY`;设置在进程内缓存。不能根据 curl 的路由推断它。注意 go.mod 的 x/net v0.58.0 代码不是这里所用的标准库 vendored 实现。 |
| DNS | `globalDNSCache.LookupHost`,dnscache v0.1.1;默认刷新窗口在容器/Kubernetes 为 30 秒,其他环境为 10 分钟,可配置。按返回地址逐个尝试 TCP,首次 TCP 成功就返回;随后 TLS/HTTP 失败不会回到这个循环尝试另一 IP。代码注释称随机选择,但该循环没有 shuffle。 |
| Linux TCP 参数 | 使用 SILO 的自定义 dialer,包含 TCP fast open/keepalive 等设置,部分参数来自 Server CLI,包括 interface、buffer、user timeout。第二轮执行了完整 Linux Server 的正常 CLI 初始化,但只验证 loopback;不能排除实际出口设备、路由或非默认 CLI 参数的交互。 |
| CA | `silo-pkg/certs.GetRootCAs` 加载系统根、Kubernetes CA 目录和 `certs/CAs`;Server 还加入自己的公开服务证书。相关 pkg CA 加载实现未在这次版本比较中变化;实际文件和路径仍可能因部署变化不同。 |
| 超时 | TCP 拨号 5 秒,TLS 握手 10 秒,响应头 1 分钟;discovery/JWKS 的 Client 没有总超时,discovery 使用的 Request 也没有调用方 context。响应体停滞不受响应头超时保护。 |
| 复用 | keep-alive 开启,idle 15 秒,TLS session cache 100。一次初始化内 discovery 与 JWKS 可复用连接;初始化重试的新 transport 没有旧连接或 session。持续首次握手失败不能用清理空闲连接解释。 |
| 请求标识 | UA 包含产品、OS、架构、模式和构建信息,产品名从 `MinIO` 变为 `Silo`;传输层禁用自动压缩。UA 规则只能在 HTTPS 被终止、HTTP 请求可见之后起作用。 |
源码:[构造与超时](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/utils.go#L651-L670)、[HTTP/TLS 参数](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/internal/http/transports.go#L44-L94)、[密码套件与曲线](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/internal/crypto/crypto.go#L52-L78)、[DNS 拨号循环](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/internal/http/dial_dnscache.go#L42-L84)、[缓存刷新](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/common-main.go#L551-L578)、[CA 加载](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/server-main.go#L383-L395)、[UA](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/update.go#L228-L266)。
比较旧 tag 与故障 tag,OpenID discovery、transport、曲线实现的变动仅为 pkg 导入路径调整;IAM 另有品牌日志变更。比较故障 tag 与本次 main,这些关键实现没有变动。因此不能把当前依赖推进当成 #154 已解决的证据。
**Console 登录阶段是另一条出站路径。** 当前及故障版 Console 通过 `GetConsoleHTTPClient` / `GlobalTransport` 再获取 discovery、交换令牌,TLS config 未指定 CurvePreferences,仍做完整证书验证;不能把“添加配置走 Server”推广为所有 Console OIDC 请求都走 Server transport。任何修复最终都必须验证完整登录。`CONSOLE_MINIO_SERVER_TLS_SKIP_VERIFY` 只针对 SILO 端点,不能作为 IdP 修复。[Console transport](https://github.com/pgsty/silo-console/blob/c103d08ec36a/api/config.go#L68-L121)、[IdP 客户端](https://github.com/pgsty/silo-console/blob/c103d08ec36a/api/tls.go#L59-L70)、[Console discovery](https://github.com/pgsty/silo-console/blob/c103d08ec36a/pkg/auth/idp/oauth2/provider.go#L428-L449)。
## 本地实验与能排除的假设
工具与原始结果放在 [issue-154/](issue-154/):
- `probe.go` 直接调用 `cmd.NewHTTPTransport()`,用同版本 `certs.GetRootCAs` 装入实验 CA;记录 TCP、TLS、HTTP 阶段。诊断参数仅修改被比较的一项。
- `fixture.go` 仅监听 IPv4 loopback;使用即时生成的 RSA 测试证书与专用 CA,提供 discovery 和有效 JWKS。正常基线限制 TLS 1.2、P-256、`TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384`,支持 HTTP/1.1 和 HTTP/2。
- `admin-check.go` 仅允许 loopback Server,验证 Console 使用的 Admin API;使用本地虚构配置和临时凭据。
- `evidence.json` 保存版本、握手参数、分阶段 trace 和完整 Server 实验结果;无客户信息、私钥或 token。
- 第二轮扩展 fixture 的合成授权码、token 和 JWKS 流程;`run-linux.py` 驱动 Linux 完整 Server 及 Console API,`linux-evidence.json` 保存十二组场景和构建身份。敏感运行值不写入结果。
### 正常服务及握手差异
三个源码版本的实际 transport 都能通过完整证书验证,以 TLS 1.2 / AES-256-GCM / HTTP/1.1 获取测试文档。**Go 1.27 本身并非不能连接 TLS 1.2、RSA 证书或这套密码套件。** 正常服务也接受诊断性的 HTTP/2。
| 构建与选项 | fixture 读到的初始握手字节数 | 支持的 group ID | 新增 ML-DSA 签名编号 |
| --- | ---: | --- | --- |
| 旧源码 + Go 1.26.5,默认 | 1497 | 4588, 29, 23, 24, 25 | 无 |
| 故障源码 + Go 1.27.1,默认 | 1509 | 同上 | 0x0904, 0x0905, 0x0906 |
| 当前源码 + Go 1.27.1,默认 | 1509 | 同上 | 同上 |
| 旧源码 + Go 1.26.5,`tlsmlkem=0` | 275 | 29, 23, 24, 25 | 无 |
| 故障/当前源码 + Go 1.27.1,`tlsmlkem=0` | 1509 | 4588, 29, 23, 24, 25 | 有 |
| 旧源码、旧依赖,只改为 Go 1.27.1 | 1509 | 4588, 29, 23, 24, 25 | 有;`tlsmlkem=0` 也不再移除 4588 |
4588 是 X25519MLKEM768,29 是 X25519,23/24/25 是 P-256/384/521。这些字节数包含本地 fixture 收到的 TLS record,测试 URL 是 IP,没有 DNS SNI;不能直接当成客户网络中的包长或 MTU 证据。默认新旧差异约 12 字节,没有证据支持“本次才突然出现巨大 ML-KEM 握手”的说法。设置上述环境选项时的差异则显著不同。
旧源码保留旧依赖、仅更换 Go 编译器后,行为随编译器变化,隔离了本次发现与 silo-pkg/Console 版本更新之间的关系。Go 1.27 官方说明也明确记录显式曲线配置不再受这些默认值开关限制,并新增 ML-DSA 支持。[Go 1.27 发布说明](https://go.dev/doc/go1.27)。本地进一步核对了两版 `crypto/tls/defaults.go`、`common.go` 的 `curvePreferences`/`supportsCurve` 和 `handshake_client.go`。
### 主动拒绝与对照结果
下列拒绝规则是人为设置的模型,**只证明机制可以产生相同症状,不证明真实 Keycloak 或其入口使用这些规则**。
| fixture 规则 | 对照结果 | 能支持的结论 |
| --- | --- | --- |
| ClientHello 带 ML-KEM 就发送 TCP RST | 旧版默认也失败;旧版 + `tlsmlkem=0` 成功;故障版 + 同选项失败;故障版显式 classical 曲线成功 | 需要旧环境选项或其他变化,才能用该机制解释升级回归。仅说 ML-KEM 不兼容不充分。 |
| ClientHello 带 ML-DSA 编号就 RST | 旧版默认成功;故障版默认/仅 classical 都失败;故障版 TLS 1.2-only 成功 | 新的签名算法列表是另一种可区分机制。ML-DSA 与 ML-KEM 不是同一项。TLS 1.2-only 会同时改变多项 ClientHello,成功不等于唯一定位 ML-DSA。 |
| 必须提供 h2 ALPN | 旧、新 Server transport 默认都失败;`-h2` 成功 | 可以解释 curl 与 Server 的不同,单独不能解释新旧版本差异。 |
| TLS 完成后,对 `Silo` UA 的 GET 发送 RST | 同一个新 transport,`MinIO` UA 成功、`Silo` UA 失败 | 报告的外层 `Get ... reset` 错误也可能来自 HTTP 层,必须先判断 TLS 是否完成。 |
| 不信任测试 CA | 返回 `*tls.CertificateVerificationError` / x509 类错误 | 与人为 RST 的错误不同;没有证据要求跳过证书验证。仍须比较客户实际连接收到的链。 |
| 正常连接复用 | 第二个请求 `reused=true`;关闭 idle 连接后重新建 TCP 可恢复 TLS session | keep-alive 与 TLS session 复用是两件事。首次新 transport 就失败时,此分支优先级低。 |
### 完整 Server、IAM 和恢复
使用三个本地编译的完整 Server,独立空数据/配置目录、独立 CA 和临时凭据;未登录真实 IdP。
| 场景 | 实际结果 |
| --- | --- |
| 旧版、故障版连接正常 IdP | discovery + JWKS 成功;一条 TLS 连接;cluster health 200;Console HTTP 200;Admin ListUsers 成功。 |
| 当前版连接持续 reset 的 IdP | IAM 持续等待;cluster health 503,`X-Minio-Server-Status: iam-offline`;Console 端口尚未提供服务;5 秒限时的 Admin ListUsers 未完成。 |
| 上述场景恢复 IdP,不重启 Server | 约 0.43 秒后 cluster/Console 200,ListUsers 成功。该时间是单次本地样本,不是恢复 SLA。 |
| 旧版 + `tlsmlkem=0`,IdP 拒绝 ML-KEM | 完整启动成功。 |
| 故障版 + `tlsmlkem=0`,相同拒绝规则 | IAM 等待、cluster 503;撤掉规则后约 1.38 秒内完整恢复,无需重启。 |
| 当前版 discovery 成功、JWKS 返回 503 | 同样阻塞 IAM;JWKS 恢复后约 0.33 秒内恢复。只验证 discovery 200 不够。 |
| 当前版先不配置 OIDC,再经 Admin API 添加 | IdP reset 时返回同型 `Get ... read tcp ... connection reset by peer`;恢复后同名创建成功,`restart=true`。失败校验没有保存该 provider。 |
在本地 IAM 阻塞的场景中,`/minio/health/live` 和 `/minio/health/ready` **仍为 200**。判断这次恢复应使用 `/minio/health/cluster` 并验证受认证操作和 Console,不能只看 ready。源码上 cluster 的 `checkHealth` 检查 IAM,而 ready 没有此检查;这是已有行为,本文不扩展为一次健康检查重构。[health 检查](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/healthcheck-handler.go#L32-L65)、[ready 检查](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/healthcheck-handler.go#L129-L186)。
现有 IAM 重试间隔随机为 0–3 秒,GET 本身还会占用网络等待时间;初始化成功前不会继续创建 IAM store,Console 又在 IAM 调用返回后启动。外部连接恢复可由现有重试自行恢复;这不意味着靠增加重试能修复持续不兼容。缺少整体请求期限、请求取消的部分是另外一个可单独修复的启动健壮性问题。[重试逻辑](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/iam.go#L342-L365)、[Console 启动顺序](https://github.com/pgsty/silo/blob/d1105bbb3d4a0afa33b3a4ac11b821235038ed0e/cmd/server-main.go#L1007-L1027)。
已通过:`CGO_ENABLED=0 GOWORK=off go test -mod=readonly -count=1 ./internal/http ./internal/config/identity/openid`。这些测试没有替代真实 Keycloak 登录或 Linux 故障网络的验证。
## 最少补充证据与诊断命令
先收集第一轮,按结果才展开后续。所有请求只取公开 discovery,不需要客户端密钥、密码、token 或私钥。不要求环境变量全量导出、配置导出或证书私钥。
### 第一轮:运行身份、网络一致性、协议阶段
在**已有故障容器/Pod 的实际网络环境**中运行;旧版也做同样检查。不要以新建默认网络容器替代。如果实际进程经启动脚本修改过环境,探针也应使用修改后的相关环境和同一 CA 挂载。
```sh
# 选择实际 Server 可执行文件;记录输出中的版本、Go、OS/架构。
silo --version
# 旧镜像的程序名可能是 minio。
# Linux:PID 设为实际 Server 进程 PID,不预设一定是 1。
# 只输出 Go 调试选项以及相关环境项是否存在,不输出代理凭据/地址。
tr '\0' '\n' < "/proc/$PID/environ" | awk '
/^GODEBUG=/ { print; next }
/^(HTTP_PROXY|HTTPS_PROXY|NO_PROXY|ALL_PROXY|http_proxy|https_proxy|no_proxy|all_proxy|SSL_CERT_FILE|SSL_CERT_DIR)=/ {
split($0, a, "="); print a[1] "=<set>"
}'
# OIDC_URL 仅在本地设为原 config_url;URL 不应带凭据或令牌。
export OIDC_URL
curl --http1.1 --connect-timeout 5 --max-time 20 -sS -o /dev/null \
-w 'ip=%{remote_ip} http=%{http_version} status=%{http_code} verify=%{ssl_verify_result}\n' "$OIDC_URL"
curl --http2 --connect-timeout 5 --max-time 20 -sS -o /dev/null \
-w 'ip=%{remote_ip} http=%{http_version} status=%{http_code} verify=%{ssl_verify_result}\n' "$OIDC_URL"
# 由维护者从对应源码构建的探针;OIDC_CA 是与 Server 相同的 CAs 目录。
./oidc-probe -ca "$OIDC_CA"
```
探针从 `OIDC_URL` 读取 URL。URL、响应体、请求头不会打印;输出的 IP 可按一致映射替换为 IP-A/IP-B。它使用真实 Server transport 构造函数,但添加 20 秒总期限、不跟随重定向、限制响应体读取为 1 MiB,并使用 `issue-154-probe` UA;因此是定位连接阶段的工具,不是完整 OIDC 流程。它没有执行 Server 的 CLI 初始化,也不继承该进程的 `--interface`、socket buffer、TCP user timeout 或存量 DNS 缓存;若使用这些非默认参数,必须对齐后才能归因。若 Server 将自己的公开服务证书也作为根信任,需要把相同公开证书加入探针的临时 CA 目录。
读结果的方法:
- `tls_start` 后 `tls_done ... err=reset`:先查 TLS ClientHello、代理 CONNECT 或 TLS 终止设备;尚不能从客户端确定 RST 由终端还是中间设备发出。
- `tls_done ... err=none`、`wrote_request` 后 reset:检查 HTTP/WAF/UA/入口路由。此时更换证书信任或密钥交换没有针对性。
- 出现 x509 类错误:比较该进程实际收到的证书链、SNI、系统 CA 和自定义 CA;保持验证开启。
- 两次 curl 的 IP 不同,或探针目标与 curl 不同:先做同 IP 对照。curl HTTP/1.1 成功也不代表 Go ClientHello 相同。
- 仅 curl h2 成功:确认实际协商的 `http_version` 是 2,再检查入口 HTTP/1.1 支持;不要先全局启用 Server HTTP/2。
### 只对命中的分支做 A/B
```sh
# TLS 阶段失败:仅移除 hybrid key exchange,仍支持 TLS 1.3 和验证证书。
./oidc-probe -ca "$OIDC_CA" -classical
# 如果旧进程确实使用 tlsmlkem=0,验证最小候选是否恢复其效果。
# 应保留原 GODEBUG 的其他相关选项;以下假定没有需要保留的其他值。
GODEBUG=tlsmlkem=0 ./oidc-probe -ca "$OIDC_CA" -default-curves
# 只有 classical 仍失败时,作为鉴别实验测试 TLS 1.2-only。
./oidc-probe -ca "$OIDC_CA" -tls12
# 只有 HTTP/ALPN 对照指向此分支时才测试。
./oidc-probe -ca "$OIDC_CA" -h2
# HTTP 阶段才失败:UA_OLD / UA_NEW 是两版实际请求的完整 UA,非密钥。
./oidc-probe -ca "$OIDC_CA" -ua "$UA_OLD"
./oidc-probe -ca "$OIDC_CA" -ua "$UA_NEW"
```
若 first-hop 使用代理,先比较两进程的代理选择和 `NO_PROXY`,不公开含密码的代理 URL。只在明确允许直连的部署中使用 `-direct`。确认直连后,可针对每个 DNS 地址运行以下两项,保留原 URL 主机名与 SNI,不能直接把 HTTPS URL 改成 IP:
```sh
curl --noproxy '*' --resolve "$OIDC_HOST:443:$IP_A" --http1.1 \
--connect-timeout 5 --max-time 20 -sS -o /dev/null \
-w 'ip=%{remote_ip} status=%{http_code} verify=%{ssl_verify_result}\n' "$OIDC_URL"
./oidc-probe -ca "$OIDC_CA" -direct -ip "$IP_A"
# 对 IP-B 重复;端口不是 443 时据实修改 --resolve。
```
仅第二次请求失败时才补 `-n 2` 与 `-n 2 -fresh`;后者重建 TCP,但同一 transport 的 TLS session cache 仍保留。若仍无法区分,下一步要的是入口侧同一时间窗口的握手失败原因/命中规则,或由客户自行脱敏后的 ClientHello 参数和 RST 阶段,不是完整认证流量包。
## 条件性最小修复方案
### 1. 路由、代理或入口策略差异已证实
优先统一有问题的 IdP 入口、修正具体域名的代理/NO_PROXY 或后端节点配置;更新错误拒绝合法 ClientHello 的 TLS 终止器/WAF。证书仍按原 hostname 和有效 CA 验证,OIDC issuer 不随意改名。若是 `Silo` UA 被规则拒绝,调整该规则;不把全产品 UA 改回 MinIO。
这是配置层处理,可以不改 SILO。工作量取决于入口归属;成功标准是原版故障二进制在原网络环境中完成 discovery、JWKS 和完整登录,而非仅 curl 200。
### 2. 证实旧环境依赖 `tlsmlkem=0`,且 classical / default-curves 对照成功
最小代码候选是为外部 OpenID 请求使用 Go 的默认曲线集合,恢复已有 GODEBUG 选项的效果;保留当前 trust pool、SNI、密码套件、超时、代理及 HTTP/1.1 行为。候选代码如下,**已在隔离副本完成上述验证,尚未应用到产品工作区**:
```go
func NewOpenIDHTTPTransport() *http.Transport {
tr := NewHTTPTransport()
tr.TLSClientConfig.CurvePreferences = nil
return tr
}
```
当时的候选只将 `cmd/iam.go` 和 `cmd/config-current.go` 两处 OpenID 调用改用该 helper,通过现有 UA wrapper 传入 `openid.LookupConfig`。后续排查确认节点互联、复制、远端存储等连接也存在同样的问题,此局部方案已被 [SILO 跨组件修复](go127-stack.md) 取代;请勿再应用下面归档的 OIDC-only 补丁。
候选已在真实 transport 探针及第二轮完整 Linux Server 中验证:`CurvePreferences=nil` + Go 1.27.1 + `tlsmlkem=0` 会移除 ML-KEM,并通过 `reject-mlkem` fixture、配置添加和合成登录。它仍不能通过 `reject-mldsa` fixture,说明此方案针对的是一个明确分支。
影响需要明确:
- 不设置 GODEBUG 时会采用完整 Go 默认集合,实测额外提供 SecP256r1MLKEM768 和 SecP384r1MLKEM1024;这也是一项兼容性变化。第二轮默认 TLS 1.3 登录通过,仍不能称为完全无行为变化或所有曲线均已覆盖。
- `GODEBUG` 是进程级设置,其他使用 Go 默认曲线的客户端也可能受到影响;本次 helper 的改动范围是 OpenID,但该环境选项本身不是逐 provider 开关。
- 禁用 hybrid 后仍有标准 ECDHE/TLS 和证书验证,失去的是相应后量子密钥交换保护。优先修正入口,临时兼容设置应有撤销条件;不能自动遇到 reset 就降级重试。
- 当前 Console IdP transport 本来使用默认曲线,同一进程中的该 GODEBUG 策略可以与之保持一致;第二轮合成登录链路已经验证,真实 IdP 和浏览器验收仍待进行。
- 隔离源码副本中的候选实现及上述本地验证已完成;真实端点 A/B 证据、生产适用性和发布是尚未完成的步骤。
### 3. 未设置上述选项,或 classical 仍失败
不要套用方案 2。若同 IP、同代理、同 UA 下仅 Go 1.27 失败,优先收集入口对新签名算法/扩展的处理证据。`-tls12` 成功只能缩小到 ClientHello/TLS 特征集合;它同时移除 hybrid key share 和 TLS 1.3 特有签名编号,不能唯一证明 ML-DSA。
优先更新入口或获得 Go 上游可复现用例。只有真实证据证明 TLS 1.2 兼容模式是必要且有效、入口又短期无法处理时,才评估**明确限定于该 IdP**的临时 TLS 1.2 选项,并保留现代 ECDHE/AEAD 和证书验证。不要把 `MaxVersion=TLS12` 全局写死,不引入自定义 ClientHello/TLS 栈或未受支持的“关闭 ML-DSA”环境选项。该分支尚不足以确定补丁或工期。
### 4. 启动健壮性可单独改进
如需小幅改善可诊断性,应单独给 discovery/JWKS 加明确阶段标识和有限总请求期限,让 IAM 重试可感知取消;在请求完成前发生 stall 时及时释放资源。日志避免输出 client secret/token,错误保留原始 cause。保留已配置 OIDC 的失败关闭行为及 IdP 恢复后的自动初始化,不自动关闭 OIDC,也不把无限重试包装成根因修复。
这类改动约 0.5–1 个工程日,可分别回归 stalled body、取消、重试后恢复;它不会使持续 RST 的 TLS 连接成功,不应替代前面鉴别。
## 临时处置与验收边界
当前可沿用报告中已经恢复服务的旧版回退状态,尽快完成上述定位。不要将旧版本长期保留视为解决方案,也不要为了它整体降级新版本依赖。本次空数据实验不证明任意生产数据/配置的降级兼容;再次切换版本前应沿既有备份和升级边界操作。
修复应按以下范围验证,均从 PGSTY 源码本地构建:
1. 对确认的分支增加最小回归:实际 OpenID transport、匹配的 ClientHello/入口拒绝条件、默认设置与明确 opt-out;确保通用和 internode transport 不变。
2. TLS 1.2(报告密码套件)与 TLS 1.3、有效自定义 CA、错误 CA 和 hostname 拒绝;适用时覆盖代理及多地址入口。不放宽证书、JWT 签名或 audience 等认证校验。
3. 完整 Server 的 discovery 与 JWKS、IdP 中断后恢复、cluster health、受认证 Admin 操作;Console 添加配置、浏览器重定向、回调、令牌交换、STS 授权和登出。mc 添加/读取配置路径应一致。
4. 在实际 Linux 容器/Pod 网络和每个 IdP 入口地址复验。证据应绑定最终 Server/Console/pkg/mc commit 和编译器;发布、镜像及生产可用性属于后续独立验收。
第二轮完成 Linux loopback 上的合成 OAuth token/STS 登录;本次没有真实 Keycloak、真实 Linux 故障路径、浏览器页面或生产数据升级测试。当前 main 对正常 fixture 成功,并不能据此宣告 #154 修复。下一项最有价值的新增证据是**实际进程的 `tlsmlkem` 设置和一个带 TLS 完成事件的同环境 GET trace**。
## 复现实验
从本任务 worktree 根目录执行;工具均使用本地源码,输出目录必须是新建临时目录。
```sh
LAB_DIR=$(mktemp -d)
CGO_ENABLED=0 GOWORK=off go build -mod=readonly -tags kqueue \
-o "$LAB_DIR/silo" .
CGO_ENABLED=0 GOWORK=off go build -mod=readonly -tags kqueue \
-o "$LAB_DIR/probe" docs/investigations/issue-154/probe.go
go build -o "$LAB_DIR/fixture" docs/investigations/issue-154/fixture.go
go build -mod=readonly -o "$LAB_DIR/admin-check" docs/investigations/issue-154/admin-check.go
"$LAB_DIR/fixture" -dir "$LAB_DIR/idp" > "$LAB_DIR/hello.jsonl" 2> "$LAB_DIR/fixture.log" &
FIXTURE_PID=$!
trap 'kill "$FIXTURE_PID" 2>/dev/null || true' EXIT
attempt=0
while [ ! -s "$LAB_DIR/idp/url" ] && [ "$attempt" -lt 50 ]; do
sleep 0.1
attempt=$((attempt + 1))
done
test -s "$LAB_DIR/idp/url" || exit 1
# 私钥仅保存在 fixture 内存。
OIDC_URL="$(cat "$LAB_DIR/idp/url")/.well-known/openid-configuration"
export OIDC_URL
"$LAB_DIR/probe" -ca "$LAB_DIR/idp/ca.pem"
printf '%s' reject-mlkem > "$LAB_DIR/idp/mode"
GODEBUG=tlsmlkem=0 "$LAB_DIR/probe" -ca "$LAB_DIR/idp/ca.pem"
GODEBUG=tlsmlkem=0 "$LAB_DIR/probe" -ca "$LAB_DIR/idp/ca.pem" -default-curves
printf '%s' reject-mldsa > "$LAB_DIR/idp/mode"
"$LAB_DIR/probe" -ca "$LAB_DIR/idp/ca.pem" -classical
"$LAB_DIR/probe" -ca "$LAB_DIR/idp/ca.pem" -tls12
kill "$FIXTURE_PID"
```
完整 Server 实验使用上述临时目录下的 `certs/CAs` 和数据目录,设置虚构的 `MINIO_IDENTITY_OPENID_CLIENT_ID`、fixture URL、临时 root 凭据,并显式指定 `--config-dir`、`--certs-dir`、loopback API/Console 地址。用 mode 文件切换 `reset`/`bad-jwks` 到 `normal`,同时检查 cluster health 和 Admin ListUsers;结果已保存在 `evidence.json`。`admin-check add` 只操作 `LAB_SERVER=127.0.0.1:port` 的实验实例,读取临时 `LAB_USER`、`LAB_PASSWORD`、`LAB_OIDC_URL`,使用虚构 OIDC secret。
跨版本构建时从各 tag `git archive` 获取源码,把 `probe.go` 放入该 module 根目录;旧版探针仅将 certs 导入替换为 `github.com/minio/pkg/v3/certs`,由旧版 go.mod 选择 PGSTY v3.11.0。分别强制 `GOTOOLCHAIN=go1.26.5` / `go1.27.1`,使用 `-mod=readonly`,不修改历史 go.mod。另将旧源码原依赖用 Go 1.27.1 构建,隔离工具链因素。若为 Linux 客户端构建探针,按已确认的架构设置 `GOOS=linux GOARCH=amd64` 或 `arm64`,不使用 Server 镜像代替。
### 第二轮 Linux 完整 Server 对照
从本任务根目录运行以下 Bash 命令。需要预先准备 Go 1.26.5、1.27.1 及源码依赖缓存,以及带 Python 3、curl/OpenSSL/HTTP2 的通用 Linux 基础镜像。`GOTOOLCHAIN` 明确选定编译器,`-mod=readonly` 保留依赖版本;实际测试构建用已安装工具链的绝对路径配合 `GOTOOLCHAIN=local`,等价地避免自动切换编译器。以下固定 `arm64` 与本次实验一致。
```bash
LAB_DIR=$(mktemp -d)
mkdir -p "$LAB_DIR/old" "$LAB_DIR/head" "$LAB_DIR/candidate" "$LAB_DIR/bin" "$LAB_DIR/out"
git archive d88f46ccee345a9c2fabe2d221d9a9e56bc11aec | tar -x -C "$LAB_DIR/old"
git archive d1105bbb3d4a0afa33b3a4ac11b821235038ed0e | tar -x -C "$LAB_DIR/head"
git archive d1105bbb3d4a0afa33b3a4ac11b821235038ed0e | tar -x -C "$LAB_DIR/candidate"
git -C "$LAB_DIR/candidate" apply "$PWD/docs/investigations/issue-154/openid-default-curves.patch"
(cd "$LAB_DIR/old" && CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 \
GOTOOLCHAIN=go1.26.5 go build -mod=readonly -tags kqueue -o "$LAB_DIR/bin/old-go126" .)
(cd "$LAB_DIR/old" && CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 \
GOTOOLCHAIN=go1.27.1 go build -mod=readonly -tags kqueue -o "$LAB_DIR/bin/old-go127" .)
(cd "$LAB_DIR/head" && CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 \
GOTOOLCHAIN=go1.27.1 go build -mod=readonly -tags kqueue -o "$LAB_DIR/bin/head-go127" .)
(cd "$LAB_DIR/candidate" && CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 \
GOTOOLCHAIN=go1.27.1 go build -mod=readonly -tags kqueue -o "$LAB_DIR/bin/candidate-go127" .)
CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 GOTOOLCHAIN=go1.27.1 \
go build -mod=readonly -o "$LAB_DIR/bin/fixture" docs/investigations/issue-154/fixture.go
CGO_ENABLED=0 GOWORK=off GOOS=linux GOARCH=arm64 GOTOOLCHAIN=go1.27.1 \
go build -mod=readonly -o "$LAB_DIR/bin/admin-check" docs/investigations/issue-154/admin-check.go
# 本机已存在的通用 Debian 12 arm64 基础镜像;不拉取镜像,不映射端口。
LINUX_BASE_IMAGE=sha256:307af7711e2e04ab75759cb42a1eef45c43c4404894c0e30dd19f742b107b922
docker run --rm --pull=never --network none \
--mount "type=bind,source=$LAB_DIR/bin,target=/lab/bin,readonly" \
--mount "type=bind,source=$LAB_DIR/out,target=/lab/out" \
--mount "type=bind,source=$PWD/docs/investigations/issue-154/run-linux.py,target=/lab/run-linux.py,readonly" \
--entrypoint python3 "$LINUX_BASE_IMAGE" /lab/run-linux.py
```
每组场景输出一行摘要,同时在独立输出目录保存 `result.json`。脚本用 `finally` 终止其 Server/fixture,容器结束自动删除。它没有导出会话 cookie、授权码、JWT、state 或临时密码;原始运行目录只用于该次隔离实验,交付证据只保留握手、结果和构建身份。
@@ -1,41 +0,0 @@
//go:build ignore
// Loopback-only lab client for the Admin API used by Console's OIDC form.
package main
import (
"context"
"fmt"
"net"
"os"
"time"
"github.com/minio/madmin-go/v3"
)
func main() {
endpoint := os.Getenv("LAB_SERVER")
host, _, err := net.SplitHostPort(endpoint)
if err != nil || host != "127.0.0.1" {
panic("LAB_SERVER must use IPv4 loopback")
}
a, err := madmin.New(endpoint, os.Getenv("LAB_USER"), os.Getenv("LAB_PASSWORD"), false)
if err != nil {
panic(err)
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if len(os.Args) > 1 && os.Args[1] == "add" {
restart, err := a.AddOrUpdateIDPConfig(ctx, "openid", "local154", "enable=on client_id=local154 client_secret=local154-placeholder config_url="+os.Getenv("LAB_OIDC_URL"), false)
fmt.Printf("add restart=%v err=%v\n", restart, err)
if err != nil {
os.Exit(1)
}
return
}
users, err := a.ListUsers(ctx)
fmt.Printf("list_users count=%d err=%v\n", len(users), err)
if err != nil {
os.Exit(1)
}
}
@@ -1,41 +0,0 @@
//go:build ignore
// Run only with the public, synthetic CA made by fixture.go.
package main
import (
"crypto/x509"
"encoding/json"
"encoding/pem"
"os"
"runtime"
"github.com/pgsty/silo-pkg/v3/certs"
)
func main() {
if len(os.Args) != 3 {
panic("usage: cert-roots synthetic-ca.pem explicit-ca-path-or-empty")
}
data, err := os.ReadFile(os.Args[1])
if err != nil {
panic(err)
}
block, _ := pem.Decode(data)
if block == nil || block.Type != "CERTIFICATE" {
panic("expected public certificate")
}
certificate, err := x509.ParseCertificate(block.Bytes)
if err != nil {
panic(err)
}
roots, err := certs.GetRootCAs(os.Args[2])
if err != nil {
panic(err)
}
_, err = certificate.Verify(x509.VerifyOptions{Roots: roots})
_ = json.NewEncoder(os.Stdout).Encode(map[string]any{
"go": runtime.Version(), "os": runtime.GOOS,
"explicit_ca": os.Args[2] != "", "trusted": err == nil,
})
}
File diff suppressed because it is too large Load Diff
-179
View File
@@ -1,179 +0,0 @@
//go:build ignore
// Loopback-only synthetic OIDC/TLS fixture. Only disposable lab identities are used.
// Rejection modes model hypotheses; they are not evidence about the user's IdP.
package main
import (
"crypto"
"crypto/rand"
"crypto/rsa"
"crypto/sha256"
"crypto/tls"
"crypto/x509"
"crypto/x509/pkix"
"encoding/base64"
"encoding/json"
"encoding/pem"
"errors"
"flag"
"fmt"
"math/big"
"net"
"net/http"
"net/url"
"os"
"path/filepath"
"slices"
"strings"
"sync"
"time"
)
type observedConn struct {
net.Conn
readBytes int
}
func (c *observedConn) Read(p []byte) (int, error) {
n, e := c.Conn.Read(p)
c.readBytes += n
return n, e
}
func (c *observedConn) reset() { _ = c.Conn.(*net.TCPConn).SetLinger(0); _ = c.Conn.Close() }
type observedListener struct{ net.Listener }
func (l observedListener) Accept() (net.Conn, error) {
c, e := l.Listener.Accept()
if e != nil {
return nil, e
}
return &observedConn{Conn: c}, nil
}
func main() {
dir := flag.String("dir", "", "isolated output directory for public CA, URL, and mode file")
tls13 := flag.Bool("tls13", false, "allow TLS 1.3 in addition to the TLS 1.2 baseline")
flag.Parse()
if *dir == "" {
panic("-dir required")
}
must(os.MkdirAll(*dir, 0700))
key, err := rsa.GenerateKey(rand.Reader, 2048)
must(err)
root := &x509.Certificate{SerialNumber: big.NewInt(1), Subject: pkix.Name{CommonName: "issue-154 local CA"}, NotBefore: time.Now().Add(-time.Hour), NotAfter: time.Now().Add(24 * time.Hour), IsCA: true, BasicConstraintsValid: true, KeyUsage: x509.KeyUsageCertSign | x509.KeyUsageDigitalSignature}
rootDER, err := x509.CreateCertificate(rand.Reader, root, root, &key.PublicKey, key)
must(err)
leaf := &x509.Certificate{SerialNumber: big.NewInt(2), Subject: pkix.Name{CommonName: "localhost"}, DNSNames: []string{"localhost"}, IPAddresses: []net.IP{net.ParseIP("127.0.0.1")}, NotBefore: root.NotBefore, NotAfter: root.NotAfter, ExtKeyUsage: []x509.ExtKeyUsage{x509.ExtKeyUsageServerAuth}, KeyUsage: x509.KeyUsageDigitalSignature}
leafDER, err := x509.CreateCertificate(rand.Reader, leaf, root, &key.PublicKey, key)
must(err)
must(os.WriteFile(filepath.Join(*dir, "ca.pem"), pem.EncodeToMemory(&pem.Block{Type: "CERTIFICATE", Bytes: rootDER}), 0600))
ln, err := net.Listen("tcp", "127.0.0.1:0")
must(err)
base := "https://" + ln.Addr().String()
must(os.WriteFile(filepath.Join(*dir, "url"), []byte(base), 0600))
mode := func() string { b, _ := os.ReadFile(filepath.Join(*dir, "mode")); return strings.TrimSpace(string(b)) }
var mu sync.Mutex
log := func(v any) { mu.Lock(); defer mu.Unlock(); _ = json.NewEncoder(os.Stdout).Encode(v) }
tc := &tls.Config{Certificates: []tls.Certificate{{Certificate: [][]byte{leafDER, rootDER}, PrivateKey: key}}, MinVersion: tls.VersionTLS12, MaxVersion: tls.VersionTLS12, CurvePreferences: []tls.CurveID{tls.CurveP256}, CipherSuites: []uint16{tls.TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384}}
if *tls13 {
tc.MaxVersion = tls.VersionTLS13
}
peerConfig := tc.Clone()
peerConfig.NextProtos = []string{"h2", "http/1.1"}
// net/http validates HTTP/2 support before GetConfigForClient; the fixture
// then deliberately selects only the reported AES-256 suite.
tc.CipherSuites = append(tc.CipherSuites, tls.TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256)
tc.GetConfigForClient = func(chi *tls.ClientHelloInfo) (*tls.Config, error) {
m := mode()
c := chi.Conn.(*observedConn)
log(map[string]any{"event": "hello", "mode": m, "bytes_read": c.readBytes, "curves": chi.SupportedCurves, "signatures": chi.SignatureSchemes, "alpn": chi.SupportedProtos, "versions": chi.SupportedVersions})
reject := m == "reset" || m == "reject-mlkem" && slices.Contains(chi.SupportedCurves, tls.CurveID(4588)) || m == "reject-mldsa" && slices.Contains(chi.SignatureSchemes, tls.SignatureScheme(0x0904)) || m == "require-h2" && !slices.Contains(chi.SupportedProtos, "h2")
if reject {
c.reset()
return nil, errors.New("synthetic ClientHello rejection")
}
return peerConfig, nil
}
server := &http.Server{TLSConfig: tc, ReadHeaderTimeout: 5 * time.Second}
var codes sync.Map
server.Handler = http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
log(map[string]any{"event": "request", "path": r.URL.Path, "protocol": r.Proto, "tls": r.TLS.Version, "cipher": r.TLS.CipherSuite, "resumed": r.TLS.DidResume, "ua": r.UserAgent()})
if mode() == "reject-silo-ua" && strings.HasPrefix(r.UserAgent(), "Silo") {
c, _, e := w.(http.Hijacker).Hijack()
if e == nil {
c.(*tls.Conn).NetConn().(*observedConn).reset()
}
return
}
w.Header().Set("Content-Type", "application/json")
switch r.URL.Path {
case "/authorize":
q := r.URL.Query()
redirect, err := url.Parse(q.Get("redirect_uri"))
if err != nil || redirect.Scheme != "http" || redirect.Hostname() != "127.0.0.1" || redirect.Path != "/oauth_callback" || q.Get("client_id") != "local154" {
http.Error(w, "loopback lab authorization only", 400)
return
}
codeBytes := make([]byte, 18)
_, err = rand.Read(codeBytes)
must(err)
code := base64.RawURLEncoding.EncodeToString(codeBytes)
codes.Store(code, q.Get("nonce"))
values := redirect.Query()
values.Set("code", code)
values.Set("state", q.Get("state"))
redirect.RawQuery = values.Encode()
http.Redirect(w, r, redirect.String(), http.StatusFound)
case "/token":
if r.Method != http.MethodPost || r.ParseForm() != nil {
http.Error(w, "bad token request", 400)
return
}
id, secret, ok := r.BasicAuth()
if !ok {
id, secret = r.Form.Get("client_id"), r.Form.Get("client_secret")
}
nonce, found := codes.LoadAndDelete(r.Form.Get("code"))
if id != "local154" || secret != "local154-placeholder" || !found || r.Form.Get("grant_type") != "authorization_code" {
w.WriteHeader(400)
_ = json.NewEncoder(w).Encode(map[string]string{"error": "invalid_grant"})
return
}
audience := id
if mode() == "bad-audience" {
audience = "different-lab-client"
}
claims, _ := json.Marshal(map[string]any{"iss": base, "sub": "local154-user", "aud": audience, "iat": time.Now().Unix(), "exp": time.Now().Add(time.Hour).Unix(), "policy": "readwrite", "nonce": nonce})
header := base64.RawURLEncoding.EncodeToString([]byte(`{"alg":"RS256","kid":"local-154","typ":"JWT"}`))
payload := header + "." + base64.RawURLEncoding.EncodeToString(claims)
hash := sha256.Sum256([]byte(payload))
signature, err := rsa.SignPKCS1v15(rand.Reader, key, crypto.SHA256, hash[:])
must(err)
if mode() == "bad-signature" {
signature[0] ^= 1
}
token := payload + "." + base64.RawURLEncoding.EncodeToString(signature)
_ = json.NewEncoder(w).Encode(map[string]any{"access_token": token, "id_token": token, "token_type": "Bearer", "expires_in": 3600})
case "/.well-known/openid-configuration":
_ = json.NewEncoder(w).Encode(map[string]any{"issuer": base, "jwks_uri": base + "/jwks", "authorization_endpoint": base + "/authorize", "token_endpoint": base + "/token", "response_types_supported": []string{"code"}, "subject_types_supported": []string{"public"}, "id_token_signing_alg_values_supported": []string{"RS256"}, "scopes_supported": []string{"openid"}})
case "/jwks":
if mode() == "bad-jwks" {
http.Error(w, "synthetic JWKS outage", 503)
return
}
_ = json.NewEncoder(w).Encode(map[string]any{"keys": []any{map[string]any{"kty": "RSA", "kid": "local-154", "use": "sig", "alg": "RS256", "n": base64.RawURLEncoding.EncodeToString(key.N.Bytes()), "e": "AQAB"}}})
default:
http.NotFound(w, r)
}
})
fmt.Fprintln(os.Stderr, base)
must(server.ServeTLS(observedListener{ln}, "", ""))
}
func must(err error) {
if err != nil {
panic(err)
}
}

Some files were not shown because too many files have changed in this diff Show More