The review found one defect shape repeated across the platform: surfaces that were declared, bound, and documented, but that nothing read. An operator configuring fullUrlRecording, bodyLogging, retry.policy, validatedDnsPinning, timeout.dns, or any of ten declared metric names got a guarantee the code never delivered. Every such surface is now in exactly one of three states -- wired for real, rejected at startup, or registered in a test-enforced gap list with its reason. No silent no-ops remain. P0: - Activate the platform from bootstrap behind app.httpclient.enabled, with a single auto-configuration importing the nine child configurations. - Give the platform a strict, repository-level ENV contract: 74 leaf fields derived from the settings record tree, unknown APP_HTTPCLIENT_* rejected. - Route typed HTTP service clients through the call kernel via KernelHttpExchangeAdapter, so they stop bypassing platform policy. - Pin dynamic-target DNS resolution to the socket for the life of a call, closing the resolve-then-connect TOCTOU / rebinding window. - Actually transmit the idempotency key, and make retry eligibility depend on transmission rather than on merely holding one. - Reject reactive authentication and reactive redirect at startup instead of declaring support that does not function. - Fix the Reactor-only Stable contract row so the lane stops failing. - Stop advertising HTTP/3 on a transport that negotiates HTTP/1. P1 covers execution and retry accounting, redirect security (per-hop target guarding, sensitive-header stripping, 303 body handling), runtime rotation and transport resource ownership keyed by generation, dynamic-target hardening (subdomain matching, global-unicast classification, strict CIDR parsing), protocol intent, pool and timeout wiring, streaming and body limits, observability parity, and OAuth single-flight refresh on a bounded pool with a bounded wait. P2 covers configuration and documentation drift, the Gradle check wiring for the four hermetic lanes, and the CI gate matrix. Two test-quality defects surfaced while closing these: the HTTP/2 stream saturation test ran against cleartext HTTP/1.1 while asserting nothing about the protocol, and an OAuth contention test slept on a latch that could fire before the callers it meant to observe. Both now assert what their names claim. Verification run: :adapter:outbound:httpclient:check and :app-bootstrap:check (checkstyle, spotless, spotbugs, and the four hermetic lanes), verifyCleanArchitectureDependencies, verifyEnvKeys, verifyOneTypePerFile, verifyDependencyLocks, the documentation and gate-matrix verifiers, and the performance lane against a real TLS+ALPN HTTP/2 server. Not executed, and tracked rather than claimed: Docker/Toxiproxy fault injection, JMH, a real QUIC/HTTP3 server, a real Spring Framework 6.2 distribution (now a delegated-pending gate), live OAuth/TLS/proxy/DNS integration, and a whole-repository check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
149 lines
8.5 KiB
Markdown
149 lines
8.5 KiB
Markdown
# Redis SDK topology lanes
|
||
|
||
These lanes exist to answer the questions the deterministic in-memory gateway cannot: how Lettuce
|
||
actually behaves during a Sentinel promotion, what a Cluster resharding does to an in-flight
|
||
command, and whether the ACL accounts grant exactly what the SDK issues.
|
||
|
||
All four have now run on Redis 7.4 and the evidence is recorded in
|
||
`docs/redis/support-matrix.md`. `.github/workflows/redis-sdk-topology.yml` runs the standalone lane
|
||
on any pull request that touches the Redis leaf, the full supported-version x topology matrix
|
||
nightly, and the same matrix on demand for a release candidate.
|
||
|
||
TLS is a lane of that matrix rather than something to wire up by hand. It is `tls`, not a
|
||
deployment mode: its shape is standalone and what it qualifies is the transport, so
|
||
`redisTopologyTest` maps the lane name to `standalone` for the tests and keeps the tag filter and
|
||
the required trust material on the lane.
|
||
|
||
## The TLS lane
|
||
|
||
`tls/compose.yml` is the standalone shape with the transport swapped. The plaintext port is turned
|
||
off entirely (`--port 0`), which is the only configuration that proves anything: a lane accepting
|
||
both would let a client that failed to negotiate TLS fall back silently and still pass.
|
||
|
||
Certificates are generated at start-up into a named volume rather than checked in — a private key
|
||
in the repository is a private key in the repository, however the file is named — and they last a
|
||
day, so a stale lane fails visibly instead of drifting.
|
||
|
||
```bash
|
||
REDIS_VERSION=7.4 docker compose -f infra/redis-sdk/tls/compose.yml up -d --wait
|
||
# The client needs the generated CA; copy it out of the volume first.
|
||
docker compose -f infra/redis-sdk/tls/compose.yml cp redis:/tls/ca.crt /tmp/redis-lane-ca.pem
|
||
cd src && ./gradlew :adapter:outbound:cache-redis:redisTopologyTest \
|
||
-Predis.topology.host=127.0.0.1 -Predis.topology.port=6390 \
|
||
-Predis.topology.mode=tls -Predis.topology.trust-material=/tmp/redis-lane-ca.pem
|
||
```
|
||
|
||
The lane refuses to run without `redis.topology.trust-material`. A TLS lane that trusts anything
|
||
qualifies nothing, so "no CA configured" is an error rather than a client with verification off.
|
||
|
||
## The ACL fixture
|
||
|
||
`acl/all-accounts.acl` provisions the accounts every lane uses. Two things about it matter, and
|
||
neither can be written in the file itself — **Redis refuses to start if an `aclfile` contains a
|
||
comment line**, so the whole file is directives and the explanation lives here.
|
||
|
||
`user default off` is the first line and is deliberate. Redis ships `default` enabled and
|
||
passwordless; while it is on, every restriction in the remaining accounts can be bypassed by simply
|
||
not authenticating, which makes the fixture decorative. Disabling it is what forces a client — and
|
||
the compose healthchecks — to pick a named account.
|
||
|
||
Every named account carries a real password — `>fixture-application`, `>fixture-advanced`, and so
|
||
on. They were `nopass`, which was the more dangerous kind of wrong: an account that accepts any
|
||
password made every assertion about authentication pass for the same reason a typo would have, so
|
||
the lane's coverage of AUTH, rotation and secret wiring was indistinguishable from no coverage.
|
||
`LiveRedisCompositionTest` now presents a wrong password on purpose and requires `WRONGPASS`, which
|
||
is only a meaningful assertion because the accounts enforce one.
|
||
|
||
The passwords are fixture values in a throwaway container and are **not** a deployment template: a
|
||
real deployment resolves each account's credential through `secret://` and never writes one into
|
||
configuration.
|
||
|
||
The accounts are also split by role, because that is how the SDK uses them. `ca-skeleton-application`
|
||
runs ordinary data commands and cannot execute a script; `ca-skeleton-application-advanced` holds
|
||
`SCRIPT LOAD` and `EVALSHA` and nothing else needs to. That separation is real rather than
|
||
decorative: `LiveRedisSemanticPortsTest` runs the rate limiter without the advanced account and
|
||
requires it to come back `Unavailable`.
|
||
|
||
## Running one
|
||
|
||
Each lane has its own endpoint, because the address a client is given is not the same kind of thing
|
||
in each topology. Standalone declares a data node; Sentinel declares a *sentinel*, from which the
|
||
primary is resolved and re-resolved when it is promoted; Cluster declares any node, from which the
|
||
rest of the topology is discovered.
|
||
|
||
```bash
|
||
# Standalone
|
||
REDIS_VERSION=7.4 docker compose -f infra/redis-sdk/standalone/compose.yml up -d --wait
|
||
cd src && ./gradlew :adapter:outbound:cache-redis:redisTopologyTest \
|
||
-Predis.topology.host=localhost -Predis.topology.port=6379 -Predis.topology.mode=standalone
|
||
|
||
# Sentinel — the port is a sentinel, and the monitored primary has to be named
|
||
REDIS_VERSION=7.4 docker compose -f infra/redis-sdk/sentinel/compose.yml up -d --wait
|
||
cd src && ./gradlew :adapter:outbound:cache-redis:redisTopologyTest \
|
||
-Predis.topology.host=localhost -Predis.topology.port=27010 \
|
||
-Predis.topology.mode=sentinel -Predis.topology.master=skeleton
|
||
|
||
# Cluster — `up --wait` waits for the `ready` gate, not just for six servers that answer PING.
|
||
# Slot assignment finishes after the nodes are healthy, and a client that connects in between sees
|
||
# CLUSTERDOWN for reasons that have nothing to do with the SDK.
|
||
REDIS_VERSION=7.4 docker compose -f infra/redis-sdk/cluster/compose.yml up -d --wait
|
||
cd src && ./gradlew :adapter:outbound:cache-redis:redisTopologyTest \
|
||
-Predis.topology.host=localhost -Predis.topology.port=7100 -Predis.topology.mode=cluster
|
||
```
|
||
|
||
Tear a lane down with `docker compose -f infra/redis-sdk/<lane>/compose.yml down -v`.
|
||
|
||
| Lane | Ports | Notes |
|
||
| --- | --- | --- |
|
||
| standalone | 6379 | bridge network, published port |
|
||
| sentinel | primary 7010, replica 7011, sentinels 27010–27012 | host network |
|
||
| cluster | nodes 7100–7105, bus 17100–17105 | host network; `ready` gates on `cluster_state:ok` |
|
||
| tls | 6390 | published port, no plaintext port at all; CA generated per run |
|
||
|
||
## Why the Sentinel and Cluster lanes use host networking
|
||
|
||
Neither topology proxies. Sentinel answers `SENTINEL get-master-addr-by-name` with the address it
|
||
monitors and the client dials that itself; a cluster client reads `CLUSTER SHARDS` and connects to
|
||
every node it names. On a bridge network those are container-internal addresses, so a client on the
|
||
host resolves a topology it cannot reach — and after a promotion it resolves a *different* one it
|
||
also cannot reach. Sharing the host network namespace makes the address the topology advertises the
|
||
address the client can use, which is the difference between testing the SDK and testing Docker's
|
||
network.
|
||
|
||
That is also why their ports are fixed rather than parameterised: the addresses are written into
|
||
Sentinel's and the cluster's own configuration at creation time, and a lane whose two halves can
|
||
disagree fails for reasons that are not the SDK's.
|
||
|
||
## Selection is by lane, not by hand
|
||
|
||
`redisTopologyTest` derives its JUnit tag expression from the declared mode: `redis-topology &
|
||
lane-<mode>`. A promotion test is meaningless without sentinels and a cross-slot test is meaningless
|
||
without a cluster, but expressing that as a runtime assumption would turn "the lane was never
|
||
started" into a green skip. Selecting by tag keeps it fail-closed — what a mode cannot prove is not
|
||
selected, and what is selected must pass.
|
||
|
||
The lane also fails closed on its endpoint: selecting `redisTopologyTest` without host, port, and
|
||
mode (and `redis.topology.master` on the Sentinel lane) is an error, never a skip. A topology test
|
||
that silently passes because it did not connect is worse than no topology test.
|
||
|
||
## `min-replicas-to-write` on the Sentinel lane
|
||
|
||
The Sentinel lane sets `min-replicas-to-write 1` and `min-replicas-max-lag 1`, and this is not
|
||
incidental configuration. Without them the lane measured a promotion in which the superseded primary
|
||
kept answering `+OK` for eleven seconds after it had been replaced: **2,086 writes acknowledged to
|
||
the caller and then discarded**, with exactly one command failing. With them the same promotion lost
|
||
one write and refused 2,020 with `NOREPLICAS`, which the SDK reports as a definite, non-ambiguous
|
||
failure a caller can act on.
|
||
|
||
Any deployment where an acknowledgement is supposed to mean something has to set these. See
|
||
`docs/redis/support-matrix.md` for the full record.
|
||
|
||
## ACL accounts
|
||
|
||
`acl/` holds one file per `CommandAccess` level. They are deliberately narrower than the SDK's own
|
||
rules, so a mistake in the SDK is still refused by the server — the account is the last boundary and
|
||
a permit never widens it.
|
||
|
||
Every lane loads the same file on every data node. Accounts are enforced per node, so "they exist on
|
||
one node" is not evidence that a topology enforces them.
|