Files
clean-architecture-backend-…/docs/redis/operations.md
T
DongHyeonkaandClaude Opus 5 5f10b791d3 chore: record pre-existing uncommitted repository state
Snapshot of the in-flight state that already existed, identically, in both
this worktree and the main checkout before this session began: the initial
HTTP Client platform implementation (previously untracked), the redis-lab
removal, and the JPA / object-storage / notification integration work.

Kept separate from this session's HTTP Client review response, which lands
in the following commit, so the two bodies of work stay reviewable apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 16:48:43 +09:00

76 lines
4.2 KiB
Markdown

# Operating the Redis SDK
## What the metrics can and cannot tell you
Every observation carries the command family, the deployment mode, and latency. None carries a key,
a field, a member, or a value — not because they would be large, but because a metric dimension
built from caller data is unbounded cardinality and, for most deployments, tenant identity in a
dashboard.
That means you can answer "which command family is slow" and "which one is failing", and you cannot
answer "which key is hot" from metrics. Use the admin plane's `SLOWLOG` projection for the first
question and `MEMORY USAGE` on a specific key for the second.
## The failures worth alerting on
| Signal | What it means | What to do |
| --- | --- | --- |
| `RedisCommandRejectedException` | the SDK refused before sending | a caller exceeded a declared bound; the reason names which one |
| `RedisCrossSlotException` | a multi-key command spans slots | the keys need a shared hash tag |
| `RedisAmbiguousExecutionException` | a write may or may not have applied | reconcile; the SDK will not retry it |
| `RedisCapabilityUnavailableException` | the server lacks the feature | a capability bean was constructed by hand, or the probe result changed |
| `SentinelFailoverObserver.ambiguousWriteCount` | non-idempotent writes lost to a promotion | each one needs reconciling; the count is the workload |
| `ClusterTopologyObserver.reshardingObserved` | `ASK`/`TRYAGAIN` seen | a slot migration is in progress; latency will be uneven until it ends |
## Things the SDK will never do for you
- Retry a non-idempotent write after a timeout. `ExecutionCertainty.AMBIGUOUS_FAILURE` is reported,
not resolved.
- Follow a cross-slot multi-key command by splitting it. It is refused instead.
- Read a whole collection, stream, or index. Every read declares a bound.
- Load a Lua script or a function library at request time. Both are deployment actions.
- Send a command it cannot classify.
- Tell you that an acknowledged write was lost. See below — this one is not a limitation you can
work around in application code.
## The write loss the client cannot see
Set these on every Redis node that can ever be a primary:
```
min-replicas-to-write 1
min-replicas-max-lag 1
```
Without them a Sentinel promotion silently destroys acknowledged writes, and this is measured, not
theoretical. In `LiveRedisSentinelPromotionTest` on the 7.4 lane, Sentinel promoted the replica and
did not demote the old primary for **eleven seconds**. The client stayed connected to a primary that
had already been replaced, wrote, and was told `+OK` **2,086 times**. Every one of those writes was
discarded when the old primary resynced. Exactly one command failed.
Nothing on the client can detect this. The server answered, so the driver recorded a success, the
SDK recorded `CONFIRMED_SUCCESS`, and the caller was told the write landed. No metric here counts
it, `SentinelFailoverObserver` cannot count it, and no retry policy helps — there was no failure to
react to. A second run of the same promotion produced sixteen thousand writes, **zero** exceptions,
and the same silent loss.
With the two settings, the identical promotion lost **one** write and refused 2,020 with
`NOREPLICAS`, which the SDK reports as a definite, non-ambiguous failure the caller can act on. That
is the whole difference: an outage you can see instead of data you cannot.
The residual window is `min-replicas-max-lag` wide and cannot be closed by configuration alone. A
write that must survive a promotion under any circumstances needs `WAIT` after it, at the cost of a
round trip to the replica — decide that per write, not globally.
## Blocking work
Blocking pops and blocking stream reads run on a dedicated connection lane. If those saturate, the
symptom is blocking calls timing out while ordinary traffic is healthy — that is the lane doing its
job, not a fault. Size the blocking pool to the number of concurrent consumers, not to request rate.
## Pub/Sub
At-most-once. A subscriber that reconnects misses whatever arrived while it was gone, and there is
no replay. Durable business events belong in a stream with a consumer group, which is at-least-once
and therefore requires idempotent consumers.