mirror of
https://github.com/pgsty/minio.git
synced 2026-08-10 00:03:29 +03:00
b6d47b739c
Findings from an adversarial review (Codex, gpt-5.6-sol at max effort) of2ff594f4band4c34d2309, each independently verified before fixing: - SBOM generation: buildx attaches a provenance attestation, so every per-arch digest names an OCI index; Syft's platform default on an amd64 runner cannot resolve an arm64-only index and the step dies. Pass --platform explicitly on all four Syft calls (the two classic lanes had the same latent defect - the renamed workflow has not run yet, which is why it never fired). - Release ordering: the HEALTHCHECK survival check now runs against the pushed architecture image before the versioned and rolling multi-arch manifests are created, so a broken health config blocks their promotion; the comment now states honestly that the arch-suffixed tags are already public at that point. - Gate assertions: tar's member-argument mode exits non-zero on any missing name, which under pipefail masked a found forbidden file when exactly one of them existed; -tv prints symlinks as 'name -> target', defeating $-anchored greps; and the licenses check proved only one-of-three. Export the rootfs once and assert every required and forbidden entry individually (busybox/sh and usr/bin/mc[li] now covered), and match the image healthcheck as an exact array instead of a substring. - Probe target vs CLI-configured servers: a probe process cannot see PID 1's argv, so --url gains EnvVar MINIO_HEALTHCHECK_URL as the documented way to point the baked-in HEALTHCHECK at a server whose address/TLS comes from command-line arguments (verified end to end: server on --address :9010, env var alone turns the container healthy). Baseline regenerated for the new env token. - IPv6 zone identifiers: serialize probe URLs via url.URL.String() so [fe80::1%eth0]:9000 becomes a valid %25-escaped URL (tests added). - Boolean flags: read --json/--quiet via Bool() so --json=false is false, instead of IsSet() which treats any occurrence as true. - Docker's HEALTHCHECK timeout raised to 10s: an outer deadline equal to the probe's own 5s always SIGKILLed the probe before it could print its diagnostic line. - test-release path filter now also triggers on cmd/healthcheck-main.go and cmd/main.go, so subcommand regressions run the image gate. Not adopted: require_text's comment-insensitivity in verify-rebrand.sh (snapshot-tripwire by design, consistent with its other assertions - the semantic check lives in the CI gate now), and full staging-then-promote tag publishing (a workflow-wide redesign shared with the classic lanes, tracked as follow-up). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
123 lines
4.4 KiB
Markdown
123 lines
4.4 KiB
Markdown
# Silo Healthcheck
|
|
|
|
Silo server exposes three un-authenticated, healthcheck endpoints liveness probe and a cluster probe at `/minio/health/live` and `/minio/health/cluster` respectively.
|
|
|
|
## Native CLI probe
|
|
|
|
The `silo` binary can probe those endpoints itself, which makes health checking possible in containers that ship no shell, `curl`, or `mc`:
|
|
|
|
```
|
|
silo healthcheck [FLAGS] [live|ready|cluster|cluster-read]
|
|
```
|
|
|
|
The check name maps 1:1 onto `/minio/health/<path>`; `live` is the default. The exit code is `0` when healthy and `1` otherwise, and one diagnostic line (including the `x-minio-server-status` and quorum headers on failure) is printed for `docker inspect` to capture. The probe target is derived the same way the server derives its own listen address — `--address` / `MINIO_ADDRESS`, with HTTPS auto-detected from `public.crt` and `private.key` in `--certs-dir` — or overridden wholesale with `--url` / `MINIO_HEALTHCHECK_URL`. The environment form exists for containers with a baked-in `HEALTHCHECK`: a probe process cannot see the server's command line, so when the server's address or TLS setup comes from CLI arguments rather than the environment, set `MINIO_HEALTHCHECK_URL` (e.g. `https://127.0.0.1:9010`) to point the built-in probe at it. Certificate verification is skipped, matching the kubelet's behavior for HTTPS probes.
|
|
|
|
Use it as an image `HEALTHCHECK` (exec form, since there may be no shell; keep the outer timeout above the probe's own 5s deadline so its diagnostic line survives):
|
|
|
|
```
|
|
HEALTHCHECK --interval=30s --timeout=10s --start-period=2m --start-interval=2s --retries=3 \
|
|
CMD ["/usr/bin/silo", "healthcheck", "ready"]
|
|
```
|
|
|
|
or as a Docker Compose healthcheck:
|
|
|
|
```
|
|
healthcheck:
|
|
test: ["CMD", "/usr/bin/silo", "healthcheck", "ready"]
|
|
interval: 5s
|
|
timeout: 10s
|
|
retries: 5
|
|
```
|
|
|
|
`silo healthcheck --maintenance cluster` answers the pre-drain question documented below: exit `0` when the node can be taken down safely, exit `1` (HTTP 412) when doing so would lose HA. Keep the `cluster` checks out of per-container liveness probes — they reflect cluster-wide quorum, not this process.
|
|
|
|
## Liveness probe
|
|
|
|
This probe always responds with '200 OK'. Only fails if 'etcd' is configured and unreachable. When liveness probe fails, Kubernetes like platforms restart the container.
|
|
|
|
```
|
|
livenessProbe:
|
|
httpGet:
|
|
path: /minio/health/live
|
|
port: 9000
|
|
scheme: HTTP
|
|
initialDelaySeconds: 120
|
|
periodSeconds: 30
|
|
timeoutSeconds: 10
|
|
successThreshold: 1
|
|
failureThreshold: 3
|
|
```
|
|
|
|
## Readiness probe
|
|
|
|
This probe always responds with '200 OK'. Only fails if 'etcd' is configured and unreachable. When readiness probe fails, Kubernetes like platforms turn-off routing to the container.
|
|
|
|
```
|
|
readinessProbe:
|
|
httpGet:
|
|
path: /minio/health/ready
|
|
port: 9000
|
|
scheme: HTTP
|
|
initialDelaySeconds: 120
|
|
periodSeconds: 15
|
|
timeoutSeconds: 10
|
|
successThreshold: 1
|
|
failureThreshold: 3
|
|
```
|
|
|
|
## Cluster probe
|
|
|
|
### Cluster-writeable probe
|
|
|
|
The reply is '200 OK' if cluster has write quorum if not it returns '503 Service Unavailable'.
|
|
|
|
```
|
|
curl http://silo1:9001/minio/health/cluster
|
|
HTTP/1.1 503 Service Unavailable
|
|
Accept-Ranges: bytes
|
|
Content-Length: 0
|
|
Server: Silo
|
|
Vary: Origin
|
|
X-Amz-Bucket-Region: us-east-1
|
|
X-Minio-Write-Quorum: 3
|
|
X-Amz-Request-Id: 16239D6AB80EBECF
|
|
X-Xss-Protection: 1; mode=block
|
|
Date: Tue, 21 Jul 2020 00:36:14 GMT
|
|
```
|
|
|
|
### Cluster-readable probe
|
|
|
|
The reply is '200 OK' if cluster has read quorum if not it returns '503 Service Unavailable'.
|
|
|
|
```
|
|
curl http://silo1:9001/minio/health/cluster/read
|
|
HTTP/1.1 503 Service Unavailable
|
|
Accept-Ranges: bytes
|
|
Content-Length: 0
|
|
Server: Silo
|
|
Vary: Origin
|
|
X-Amz-Bucket-Region: us-east-1
|
|
X-Minio-Write-Quorum: 3
|
|
X-Amz-Request-Id: 16239D6AB80EBECF
|
|
X-Xss-Protection: 1; mode=block
|
|
Date: Tue, 21 Jul 2020 00:36:14 GMT
|
|
```
|
|
|
|
### Checking cluster health for maintenance
|
|
|
|
You may query the cluster probe endpoint to check if the node which received the request can be taken down for maintenance, if the server replies back '412 Precondition Failed' this means you will lose HA. '200 OK' means you are okay to proceed.
|
|
|
|
```
|
|
curl http://silo1:9001/minio/health/cluster?maintenance=true
|
|
HTTP/1.1 412 Precondition Failed
|
|
Accept-Ranges: bytes
|
|
Content-Length: 0
|
|
Server: Silo
|
|
Vary: Origin
|
|
X-Amz-Bucket-Region: us-east-1
|
|
X-Amz-Request-Id: 16239D63820C6E76
|
|
X-Xss-Protection: 1; mode=block
|
|
X-Minio-Write-Quorum: 3
|
|
Date: Tue, 21 Jul 2020 00:35:43 GMT
|
|
```
|