Files
DongHyeonkaandClaude Opus 5 5f10b791d3 chore: record pre-existing uncommitted repository state
Snapshot of the in-flight state that already existed, identically, in both
this worktree and the main checkout before this session began: the initial
HTTP Client platform implementation (previously untracked), the redis-lab
removal, and the JPA / object-storage / notification integration work.

Kept separate from this session's HTTP Client review response, which lands
in the following commit, so the two bodies of work stay reviewable apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 16:48:43 +09:00

4.2 KiB

Operating the Redis SDK

What the metrics can and cannot tell you

Every observation carries the command family, the deployment mode, and latency. None carries a key, a field, a member, or a value — not because they would be large, but because a metric dimension built from caller data is unbounded cardinality and, for most deployments, tenant identity in a dashboard.

That means you can answer "which command family is slow" and "which one is failing", and you cannot answer "which key is hot" from metrics. Use the admin plane's SLOWLOG projection for the first question and MEMORY USAGE on a specific key for the second.

The failures worth alerting on

Signal What it means What to do
RedisCommandRejectedException the SDK refused before sending a caller exceeded a declared bound; the reason names which one
RedisCrossSlotException a multi-key command spans slots the keys need a shared hash tag
RedisAmbiguousExecutionException a write may or may not have applied reconcile; the SDK will not retry it
RedisCapabilityUnavailableException the server lacks the feature a capability bean was constructed by hand, or the probe result changed
SentinelFailoverObserver.ambiguousWriteCount non-idempotent writes lost to a promotion each one needs reconciling; the count is the workload
ClusterTopologyObserver.reshardingObserved ASK/TRYAGAIN seen a slot migration is in progress; latency will be uneven until it ends

Things the SDK will never do for you

  • Retry a non-idempotent write after a timeout. ExecutionCertainty.AMBIGUOUS_FAILURE is reported, not resolved.
  • Follow a cross-slot multi-key command by splitting it. It is refused instead.
  • Read a whole collection, stream, or index. Every read declares a bound.
  • Load a Lua script or a function library at request time. Both are deployment actions.
  • Send a command it cannot classify.
  • Tell you that an acknowledged write was lost. See below — this one is not a limitation you can work around in application code.

The write loss the client cannot see

Set these on every Redis node that can ever be a primary:

min-replicas-to-write 1
min-replicas-max-lag 1

Without them a Sentinel promotion silently destroys acknowledged writes, and this is measured, not theoretical. In LiveRedisSentinelPromotionTest on the 7.4 lane, Sentinel promoted the replica and did not demote the old primary for eleven seconds. The client stayed connected to a primary that had already been replaced, wrote, and was told +OK 2,086 times. Every one of those writes was discarded when the old primary resynced. Exactly one command failed.

Nothing on the client can detect this. The server answered, so the driver recorded a success, the SDK recorded CONFIRMED_SUCCESS, and the caller was told the write landed. No metric here counts it, SentinelFailoverObserver cannot count it, and no retry policy helps — there was no failure to react to. A second run of the same promotion produced sixteen thousand writes, zero exceptions, and the same silent loss.

With the two settings, the identical promotion lost one write and refused 2,020 with NOREPLICAS, which the SDK reports as a definite, non-ambiguous failure the caller can act on. That is the whole difference: an outage you can see instead of data you cannot.

The residual window is min-replicas-max-lag wide and cannot be closed by configuration alone. A write that must survive a promotion under any circumstances needs WAIT after it, at the cost of a round trip to the replica — decide that per write, not globally.

Blocking work

Blocking pops and blocking stream reads run on a dedicated connection lane. If those saturate, the symptom is blocking calls timing out while ordinary traffic is healthy — that is the lane doing its job, not a fault. Size the blocking pool to the number of concurrent consumers, not to request rate.

Pub/Sub

At-most-once. A subscriber that reconnects misses whatever arrived while it was gone, and there is no replay. Durable business events belong in a stream with a consumer group, which is at-least-once and therefore requires idempotent consumers.