7100 lines
259 KiB
Markdown
7100 lines
259 KiB
Markdown
# Redis Production Capability Deep Design
|
|
|
|
- Date: 2026-07-26
|
|
- Status: 상세 설계 완료, 5개 standalone `implemented-candidate`, selected/R2 없음
|
|
- Scope: Redis 전용 production capability와 단계적 구현 설계
|
|
- Baseline: Java 21, Spring Boot 4.0.0, Gradle multi-module Clean Architecture template
|
|
- Parent:
|
|
[Production Capability Platform Design](2026-07-26-production-capability-platform-design.md)
|
|
|
|
## 0. 구현 상태
|
|
|
|
2026-07-30 기준 구현된 범위:
|
|
|
|
- `application-core`의 provider-neutral `CacheRegionPort`와 hit/negative/miss/schema/unavailable
|
|
결과 구분;
|
|
- source revision과 application consistency intent를 담는 record/invalidation outcome;
|
|
- namespace, key/hash version, single hash slot, bounded digest를 고정하는 `RedisKeyBuilder`;
|
|
- opaque ID용 SHA-256 및 민감한 composite scope용 length-prefixed HMAC-SHA-256;
|
|
- `compare-and-delete-v1`, `compare-and-expire-v1`, `set-if-absent-with-ttl-v1` Lua resource;
|
|
- exact SHA-256와 status/signature를 기록한 R0 `program-set.json`;
|
|
- generic application API가 아닌 package-private `RedisAtomicPrimitives` internal R0 foundation과
|
|
compatibility failure;
|
|
- managed Lettuce standalone connection lifecycle과 finite command timeout;
|
|
- `EVALSHA` 우선, 정확한 `NOSCRIPT`에만 catalog script를 `SCRIPT LOAD`하고 digest를 검증한 뒤
|
|
`EVALSHA`를 한 번 재시도하는 production executor;
|
|
- versioned digest-protected bounded binary cache envelope, positive/negative TTL, invalidate와
|
|
corrupt/future/unavailable 구분을 제공하는 `CacheRegionPort<String,String>` reference adapter;
|
|
- envelope v2의 absolute soft/hard expiry, injected clock freshness 판정, deterministic
|
|
policy-revision/key jitter, hard minimum과 physical Redis TTL 일치;
|
|
- framework-free `CacheAsideExecutor`와 typed source/result/cancellation contract;
|
|
- maximum in-flight key/waiter/source concurrency/admission/load deadline을 제한하는 local
|
|
single-flight와 source bulkhead, abandoned-flight opportunistic reaping;
|
|
- authoritative absence만 negative-cache하고 classified transient failure에만 hard-expiry 전
|
|
stale fallback을 허용하는 application policy;
|
|
- HMAC key secret/namespace/value bound typed settings와 disabled zero-connection composition;
|
|
- `managed`/`external` client mode를 통한 결정적 runtime 선택;
|
|
- reconnect command replay 차단, finite Lettuce request queue와 client-side admission;
|
|
- bounded Lua `GETRANGE` read로 wire bulk reply를 envelope maximum + 1 byte로 제한하고
|
|
oversized 외부 value를 typed incompatible schema로 격리;
|
|
- managed runtime 활성화 시 Redis host 누락을 `localhost`로 숨기지 않는 startup fail-fast;
|
|
- generic Lua executor/descriptor와 raw-key typed primitive를 package-private collaborator로
|
|
닫고 Spring composition에는 semantic cache port만 노출;
|
|
- 명시적 Redis 7.2/7.4 standalone service lane의 실제 TTL expiry, compare-delete Lua,
|
|
oversized bulk-reply 차단 검증;
|
|
- `shared-contract`의 provider-neutral edge rate-limit request/policy/decision/outcome/port;
|
|
- fixed window, sliding-window counter, token bucket의 versioned one-key Lua와 bounded
|
|
structured MULTI reply parser;
|
|
- private HMAC key, Redis server time, clock regression clamp, denial-no-consume, finite state
|
|
TTL과 pre-send/post-dispatch failure certainty를 보존하는 semantic provider;
|
|
- cache와 endpoint/connection/admission/settings를 공유하지 않는 coordination-role 전용
|
|
`app.rate-limit` composition과 disabled zero-side-effect gating;
|
|
- 세 알고리즘을 실제 standalone Redis에 실행하도록 선택 가능한 service qualification lane;
|
|
- request-replay idempotency, cache refresh soft lease, versioned session repository semantic
|
|
provider와 각 card-owned standalone/security/fault/compatibility evidence;
|
|
- cache generation/revision invalidation, bounded local L1, authenticated invalidation hint,
|
|
semantic health/metrics와 standalone TLS+named ACL evidence.
|
|
|
|
아직 구현되지 않은 범위:
|
|
|
|
- refresh-ahead와 probabilistic early refresh;
|
|
- Redis Functions 배포와 program upgrade/rollback compatibility matrix;
|
|
- Sentinel runtime, Cluster production qualification, k3s/multi-node/failover/rotation,
|
|
effective eviction/persistence attestation;
|
|
- fenced coordination과 multi-process/pod session 및 L1/L2 distributed qualification;
|
|
- Phase 1의 전체 acceptance와 R2/R3 승격 증거.
|
|
|
|
현재 registry의 cache, edge rate limit, request-replay idempotency, cache refresh soft lease,
|
|
session card는 standalone promotion topology의 `implemented-candidate`다. fenced coordination만
|
|
`not-implemented`다. `implemented-candidate`는 구현과 card-owned evidence lane을 뜻할 뿐 release
|
|
selection이나 R2 qualification이 아니다. checked-in `selected` card가 0개이므로 Redis capability
|
|
전체 또는 어떤 production topology에도 R2 release claim을 하지 않는다.
|
|
|
|
## 1. 설계 판정
|
|
|
|
설계 착수 당시 `adapter:outbound:cache-redis`는 실제 Redis client, connection, topology, TTL,
|
|
codec, atomic program, failure semantics가 없는 R0 extension seam이었다. 2026-07-30 현재 위 5개
|
|
semantic provider는 standalone `implemented-candidate`이며 standalone TLS+named ACL과 bounded
|
|
fault evidence도 있다. 그러나 selection, Sentinel/Cluster, multi-node/failover/rotation,
|
|
effective eviction/persistence attestation과 R3 증거가 없으므로 production-ready/R2라는 단일
|
|
label을 붙이지 않는다.
|
|
|
|
이번 설계는 다음 구조를 선택한다.
|
|
|
|
1. 정확히 19개인 현재 leaf registry는 우선 유지한다.
|
|
2. 물리 모듈 `adapter:outbound:cache-redis`는 R2 단계에서 Redis technology capability
|
|
provider로 확장한다.
|
|
3. application/domain에는 범용 `RedisPort`, raw key, command, Lua, `RedisTemplate`, Lettuce,
|
|
Spring Session 타입을 노출하지 않는다.
|
|
4. cache, edge rate limit, idempotency, efficiency lease, fenced coordination, session은 서로
|
|
다른 semantic contract와 failure policy를 가진다.
|
|
5. 최소한 `cache`, `coordination`, `session` Redis role을 서로 다른 deployment로 격리한다.
|
|
6. Redis의 command 직렬 실행, Lua/Function atomicity, `WAIT`, AOF, Sentinel, Cluster를
|
|
cross-store exactly-once나 strong correctness lock으로 표현하지 않는다.
|
|
7. 안전한 Redis operation은 versioned program catalog로 제공하되 application이 임의 command나
|
|
script를 실행하게 하지 않는다.
|
|
8. 운영 profile은 real client, bounded resources, TLS/ACL, health, metrics, failure injection,
|
|
topology test, runbook까지 갖춰야 R2/R3로 표시한다.
|
|
|
|
현재 상태와 목표는 다음과 같다.
|
|
|
|
| Capability | 현재 | 목표 |
|
|
| --- | --- | --- |
|
|
| Redis runtime | canonical role router와 managed/external Lettuce runtime, standalone candidate | Sentinel runtime과 Cluster production qualification |
|
|
| Cache | standalone `implemented-candidate`; generation/soft lease/bounded L1과 TLS/ACL/fault lane | multi-process L1/L2와 HA/persistence/eviction attestation |
|
|
| Rate limit | fixed/sliding-counter/token-bucket standalone `implemented-candidate` | HA topology, failover와 R3 evidence |
|
|
| Idempotency | owner-safe Redis V2 standalone `implemented-candidate`; JDBC provider와 명시적 선택 | actual-used image/event evidence와 selected promotion |
|
|
| Lock | Redis efficiency lease candidate; fenced coordination은 `not-implemented` | protected-resource stale fencing-token rejection |
|
|
| Session | JWT와 isolated Redis Session profile; Redis는 standalone `implemented-candidate` | multi-process/pod와 failover/rotation qualification |
|
|
| Atomic helper | versioned closed Lua catalog와 typed internal facade | Redis Functions upgrade/rollback matrix |
|
|
| Topology | standalone candidate; Cluster code seam; Sentinel runtime 미구현 | Sentinel/Cluster/k3s multi-node qualification |
|
|
| Failure | capability별 typed degraded/unavailable/indeterminate와 bounded fault lane | 실제 topology event chain과 persistence/restart evidence |
|
|
| CI | strict registry matrix, real candidate lanes, sanitized artifact/reconciler | actual-used image attestation과 actual fault-event capture |
|
|
|
|
설계가 완료되었다는 뜻은 구현 계약과 단계가 결정되었다는 뜻이다. 현재 Redis runtime이
|
|
production-ready가 되었다는 뜻은 아니다.
|
|
|
|
## 2. 기존 통합 설계와 이번 심화 설계의 관계
|
|
|
|
상위 통합 설계는 다음 결정을 이미 내렸다.
|
|
|
|
- semantic port와 provider를 분리한다.
|
|
- cache, coordination, session Redis role을 격리한다.
|
|
- cache fail-open을 correctness capability에 재사용하지 않는다.
|
|
- rate-limit algorithm을 policy별로 선택한다.
|
|
- idempotency와 lock에 owner token과 fencing을 도입한다.
|
|
- Redis SDK나 raw command를 core에 노출하지 않는다.
|
|
|
|
이번 문서는 그 결정을 실제 구현자가 임의로 해석하지 않도록 다음을 추가로 고정한다.
|
|
|
|
- 현재 코드의 정확한 결함과 제거 순서;
|
|
- physical module 유지와 향후 split 조건;
|
|
- application/shared/web/bootstrap의 소유권;
|
|
- role, deployment, connection, key, codec, program의 구체 계약;
|
|
- cache lookup/write/invalidation 결과 모델;
|
|
- cache-aside, negative cache, stale, refresh, L1/L2, invalidation 전략;
|
|
- rate-limit algorithm별 상태, 비용, 원자성, fallback;
|
|
- lease, renewal, lost state, fencing, unknown outcome;
|
|
- Redis idempotency state machine과 cross-store 한계;
|
|
- Spring Session profile, serializer, expiry, concurrent mutation, logout;
|
|
- standalone/Sentinel/Cluster, replication, persistence, eviction의 실제 보장;
|
|
- client reconnect/replay, timeout, queue, pool, backpressure;
|
|
- security, observability, health, graceful shutdown, runbook;
|
|
- real-service, concurrency, failover, memory, compatibility CI.
|
|
|
|
세부 내용이 상위 문서의 Redis 요약과 다를 경우 이 Redis 전용 문서가 Redis 범위의 정본이다.
|
|
상위 문서의 다른 capability 결정은 변경하지 않는다.
|
|
|
|
### 2.1 Normative decision ledger
|
|
|
|
긴 문서에서 결정을 다시 추론하지 않도록 구현과 리뷰는 다음 정본 위치를 사용한다.
|
|
|
|
| 결정 | 정본 |
|
|
| --- | --- |
|
|
| capability/provider 소유권과 모듈 경계 | §7–§8 |
|
|
| readiness/guarantee 용어 | §9 |
|
|
| role/deployment 격리 | §10, §27–§29 |
|
|
| key/codec/program manifest | §11–§13 |
|
|
| common primitive와 자료구조 안전 기준 | §14 |
|
|
| cache 계약과 source 결과 | §15–§18 |
|
|
| rate-limit 알고리즘과 결과 | §19–§21 |
|
|
| lease/fencing과 idempotency | §22–§24 |
|
|
| Redis Session 보안 계약 | §25 |
|
|
| client/replay/timeout | §31 |
|
|
| 설정·activation SSOT | §32–§33 |
|
|
| secret material과 ACL | §34 |
|
|
| CI task/lane/evidence | §37 |
|
|
| dependency ownership | §38 |
|
|
| 단계별 readiness 승격 | §39–§40 |
|
|
|
|
표의 링크 대상보다 예시 YAML이나 migration alias가 우선하지 않는다. 상충하는 두 설정이
|
|
존재하면 임의 precedence를 선택하지 않고 startup을 실패시킨다.
|
|
|
|
## 3. 설계 착수 당시 증거 기반 baseline
|
|
|
|
이 절 전체는 구현 전 repository를 조사한 2026-07-26 역사적 baseline이다. 아래의 “현재”는 그
|
|
조사 시점을 가리키며 2026-07-30 구현 상태를 설명하지 않는다. 최신 구현/readiness truth는 §0,
|
|
§1의 현재 열, checked-in `src/config/redis/readiness-cards.yaml`, Redis leaf README를 따른다.
|
|
|
|
### 3.1 실제 Redis client가 없다
|
|
|
|
현재 leaf의 production dependency는 다음뿐이다.
|
|
|
|
```text
|
|
shared-contract
|
|
adapter:outbound:support
|
|
spring-boot-autoconfigure
|
|
slf4j-api
|
|
```
|
|
|
|
Spring Data Redis, Lettuce, Jedis, Redisson 중 어떤 runtime도 없다.
|
|
|
|
`RedisCacheAdapterConfig`는 `app.cache.redis.enabled=true`이면 `RedisClient` bean을 요구하지만,
|
|
production 구현은 없다. 테스트가 anonymous fake를 주입해서 bean gating만 확인한다. 따라서 현재
|
|
enable flag는 “Redis가 동작한다”가 아니라 “forking project가 별도 client를 구현했을 때 seam을
|
|
기여한다”는 뜻이다.
|
|
|
|
### 3.2 application이 소비할 합법적인 port가 없다
|
|
|
|
`CacheStore`, `CacheBackend`, `CacheStoreRouter`는 모두 outbound adapter 내부 타입이다.
|
|
`application-core`는 adapter leaf에 의존할 수 없으므로 use case가 이 router를 합법적으로 주입받을
|
|
수 없다. production consumer 검색 결과도 0개이며 bootstrap test만 router를 사용한다.
|
|
|
|
이는 cache code가 존재하지만 Clean Architecture의 실제 outbound port가 존재하지 않는 상태다.
|
|
|
|
### 3.3 logical region이 physical key에 반영되지 않는다
|
|
|
|
현재 router는 `logicalName`으로 backend만 선택하고 backend에는 raw `key`만 전달한다.
|
|
|
|
```text
|
|
router.get("worklog", "42") -> redis.get("42")
|
|
router.get("codes", "42") -> redis.get("42")
|
|
```
|
|
|
|
두 region이 같은 backend를 사용하면 충돌한다. application, environment, tenant, capability,
|
|
region, key schema version도 구분되지 않는다. session이나 idempotency를 같은 backend에
|
|
연결한다면 더 치명적이다.
|
|
|
|
### 3.4 cache lifecycle을 표현할 수 없다
|
|
|
|
현재 계약에는 다음이 없다.
|
|
|
|
- positive/negative TTL;
|
|
- soft/hard TTL;
|
|
- TTL jitter;
|
|
- invalidate/delete;
|
|
- conditional write;
|
|
- bulk get/evict;
|
|
- namespace generation;
|
|
- schema version;
|
|
- payload size;
|
|
- corruption outcome;
|
|
- after-commit invalidation;
|
|
- stampede suppression;
|
|
- stale-if-error.
|
|
|
|
특히 `put(key, value)`에 TTL이 없으므로 단순 `SET` 구현은 immortal cache를 만든다.
|
|
|
|
### 3.5 장애와 miss가 합쳐진다
|
|
|
|
`FailOpenCacheStore`는 모든 backend를 중앙에서 감싸고 모든 `Exception`을 다음처럼 처리한다.
|
|
|
|
```text
|
|
get failure -> Optional.empty()
|
|
put failure -> swallow
|
|
```
|
|
|
|
따라서 정상 miss, timeout, connection failure, wrong-type, corrupt payload, codec bug,
|
|
programming defect를 caller가 구분할 수 없다. 성능 최적화용 cache의 일부 장애에는 fail-open이
|
|
가능하지만, codec bug까지 miss로 숨기는 것은 장애 증폭과 source overload를 만든다. session,
|
|
idempotency, strict rate limit, lock에 이 decorator를 재사용하는 것은 금지한다.
|
|
|
|
### 3.6 multi-instance mode는 현재 조립할 수 없다
|
|
|
|
`APP_MULTI_INSTANCE_ENABLED=true`는 bean name으로 다음 다섯 개를 요구한다.
|
|
|
|
```text
|
|
distributedLockProvider
|
|
cacheStampedeProtection
|
|
outboxLeaderElection
|
|
distributedRateLimiter
|
|
migrationStartupRunner
|
|
```
|
|
|
|
현재 production composition에는 `cacheStampedeProtection`과 `distributedRateLimiter` 두 bean이
|
|
없고 test configuration만 다섯 이름의 plain `Object`를 제공한다. 그러므로 현재 실제 composition은
|
|
multi-instance mode에서 반드시 startup failure가 난다. 더 큰 문제는 단순 bean name 검사가
|
|
provider topology나 guarantee를 검증하지 않는다는 점이다.
|
|
|
|
### 3.7 rate limit은 fixed-window local map 하나다
|
|
|
|
현재 `RateLimitAlgorithm`은 `FIXED_WINDOW`만 제공한다. `FixedWindowRateLimiter`는 process-local
|
|
`ConcurrentHashMap`을 사용하므로 pod마다 quota가 따로 존재하고 key removal policy가 없다.
|
|
|
|
추가 결함은 다음과 같다.
|
|
|
|
- authenticated key에 route/policy dimension이 없다.
|
|
- raw principal/IP를 Redis key로 옮길 위험이 있다.
|
|
- window boundary concurrency contract가 없다.
|
|
- `Retry-After`가 decision이 아니라 고정 1초에 묶여 있다.
|
|
- invalid limit/window가 fail-fast하지 않고 default로 조용히 바뀐다.
|
|
- key cardinality budget과 cleanup이 없다.
|
|
|
|
### 3.8 idempotency와 lock port는 Redis provider를 안전하게 수용하지 못한다
|
|
|
|
현재 `IdempotencyStorePort`는:
|
|
|
|
```text
|
|
find(scope)
|
|
tryBegin(scope, fingerprint, expiresAt)
|
|
complete(scope, response)
|
|
discard(scope)
|
|
```
|
|
|
|
형태다. `complete`와 `discard`에 owner token이 없어 old owner의 lease가 만료된 후 new owner가
|
|
claim해도 stale owner가 새 record를 덮어쓰거나 삭제할 수 있다. processing lease와 completed
|
|
response replay TTL도 하나의 TTL로 합쳐져 있다.
|
|
|
|
현재 `DistributedLockPort`는 acquired handle의 `close()`만 제공한다. owner token, renewal,
|
|
lease-lost, current validity, fencing token, acquire/release의 unknown outcome을 표현하지 못한다.
|
|
문서상 efficiency lock인 점은 올바르지만 Redis provider를 붙일 계약으로는 불충분하다.
|
|
|
|
### 3.9 configuration registry와 runtime이 어긋난다
|
|
|
|
registry에는 Redis host, port, password, cache TTL 관련 key가 있지만 typed settings와 실제
|
|
client 소비자는 없다. `application.yml`에는 사실상 enabled flag만 있다.
|
|
|
|
빠진 운영 설정은 다음과 같다.
|
|
|
|
- standalone/Sentinel/Cluster;
|
|
- endpoint discovery와 DB index;
|
|
- TLS와 hostname verification;
|
|
- ACL username/password secret reference;
|
|
- connect/command/acquire/overall timeout;
|
|
- request queue와 pool bound;
|
|
- topology refresh와 redirect limit;
|
|
- primary/replica read policy;
|
|
- client name;
|
|
- shutdown/drain;
|
|
- role별 required/readiness;
|
|
- server/program/schema compatibility.
|
|
|
|
### 3.10 현재 test가 증명하는 범위
|
|
|
|
현행 focused test는 통과한다.
|
|
|
|
```text
|
|
./gradlew :adapter:outbound:cache-redis:test --rerun-tasks --console=plain
|
|
19 tests, failures/errors/skipped 0
|
|
```
|
|
|
|
이 test는 fake seam의 routing과 fail-open 동작을 증명한다. 실제 Redis command, TTL, Lua,
|
|
eviction, replication, failover, Cluster slot, TLS/ACL, session, concurrency는 증명하지 않는다.
|
|
|
|
## 4. 범위와 명시적 비범위
|
|
|
|
### 4.1 R2 baseline에 포함
|
|
|
|
- Spring Data Redis + Lettuce real provider;
|
|
- standalone과 managed/Sentinel topology의 production profile;
|
|
- Cluster-compatible key/program 설계;
|
|
- cache, coordination, session role 분리;
|
|
- typed settings와 startup validation;
|
|
- versioned key builder와 codec;
|
|
- versioned Function/Lua program catalog;
|
|
- cache-aside, negative cache, TTL jitter, invalidate;
|
|
- soft/hard TTL과 stale-if-error;
|
|
- local single-flight와 optional Redis refresh lease;
|
|
- fixed window, sliding counter, token bucket rate limiter;
|
|
- owner-safe Redis efficiency lease;
|
|
- owner-safe Redis idempotency request replay profile;
|
|
- Redis-backed Spring Session profile;
|
|
- TLS, ACL, secret rotation contract;
|
|
- capability metrics, health, traces, logs;
|
|
- real Redis, concurrency, memory, failure integration test;
|
|
- explicit provider selection and no side effect when disabled.
|
|
|
|
### 4.2 R2에서 열어둘 advanced operation
|
|
|
|
- exact sliding-window log;
|
|
- GCRA;
|
|
- L1 local cache + Redis L2;
|
|
- probabilistic early refresh;
|
|
- refresh-ahead;
|
|
- client-side tracking;
|
|
- Pub/Sub cache invalidation hint;
|
|
- namespace generation invalidation;
|
|
- fenced lease;
|
|
- bounded distributed semaphore;
|
|
- leader election;
|
|
- indexed Spring Session repository;
|
|
- Redis Function provisioning mode;
|
|
- `WAIT`/`WAITAOF` acknowledgement profile;
|
|
- replica reads for explicitly stale-tolerant cache;
|
|
- compression;
|
|
- multi-region cache warming.
|
|
|
|
각 advanced operation은 enable만으로 R2가 되지 않는다. 별도 capability card와 contract test가
|
|
필요하다.
|
|
|
|
### 4.3 R3에서 검증할 항목
|
|
|
|
- Redis Cluster reshard와 rolling topology change;
|
|
- Sentinel/Cluster failover under partition;
|
|
- serializer/key/program rolling compatibility;
|
|
- supported Redis version matrix;
|
|
- credential/certificate rotation without global outage;
|
|
- persistence recovery와 declared RPO 검증;
|
|
- capacity/latency soak;
|
|
- fenced consumer의 실제 stale-token rejection;
|
|
- multi-region topology와 region failover.
|
|
|
|
### 4.4 비범위
|
|
|
|
- 모든 future domain을 위한 universal repository;
|
|
- raw Redis data-structure facade를 application에 제공;
|
|
- 임의 Lua/Function source 실행 API;
|
|
- Redis를 authoritative relational database처럼 사용;
|
|
- generic cross-store transaction;
|
|
- exactly-once side effect 보장;
|
|
- Redis lock만으로 business invariant 보장;
|
|
- synchronous HTTP leaky-bucket queue;
|
|
- Redis Pub/Sub을 durable business event bus로 사용;
|
|
- Redis Streams를 현재 messaging leaf에 암묵적으로 추가;
|
|
- 실제 workload 없이 단일 maxmemory, pool size, timeout, TPS를 정답으로 고정;
|
|
- Redis server의 deployment IaC 전체 구현.
|
|
|
|
## 5. HARD invariants
|
|
|
|
구현은 다음 조건을 모두 지켜야 한다.
|
|
|
|
1. `domain-core`에는 Redis, cache, session, rate-limit 기술 개념이 없다.
|
|
2. `application-core`에는 Spring, Lettuce, Redis command, key syntax, Lua, serializer SDK가 없다.
|
|
3. application use case는 `RedisTemplate`, connection, raw command executor를 받지 않는다.
|
|
4. transport rate limit과 business quota를 같은 port로 합치지 않는다.
|
|
5. session repository를 application port로 추상화하지 않는다.
|
|
6. cache miss와 backend unavailable을 같은 결과로 합치지 않는다.
|
|
7. codec/schema/programming 오류는 fail-open miss로 숨기지 않는다.
|
|
8. correctness capability는 evictable cache role에 bind하지 않는다.
|
|
9. Redis database number와 prefix를 workload isolation으로 간주하지 않는다.
|
|
10. expirable write는 value write와 TTL을 한 atomic command/program에서 수행한다.
|
|
11. lock release/renew는 owner token을 비교한다. blind `DEL`/`PEXPIRE`는 금지한다.
|
|
12. idempotency complete/release는 owner token과 claim revision을 비교한다.
|
|
13. multi-command read/decide/write를 “Redis가 single-thread이므로 안전”하다고 설명하지 않는다.
|
|
14. Lua/Function은 bounded complexity와 bounded state growth를 가져야 한다.
|
|
15. program은 classpath/provisioned artifact로 version/checksum이 고정된다.
|
|
16. runtime caller가 동적 script source나 key name을 programmatically 생성하지 않는다.
|
|
17. Cluster multi-key atomic operation은 같은 slot임을 key builder와 test가 보장한다.
|
|
18. timeout/reset 후 mutation 결과를 자동으로 `FAILED`라고 단정하지 않는다.
|
|
19. non-idempotent mutation을 결과 확인 없이 무조건 retry하지 않는다.
|
|
20. `WAIT`, `WAITAOF`, AOF, replica를 strong consistency나 zero-loss로 표현하지 않는다.
|
|
21. Redis lease를 fencing 없는 correctness lock으로 표현하지 않는다.
|
|
22. Pub/Sub/keyspace notification을 durable invalidation이나 expiry source of truth로 사용하지 않는다.
|
|
23. cache DB update와 Redis update가 atomic하다고 표현하지 않는다.
|
|
24. raw PII, credential, token, session ID, idempotency key를 Redis key/log/metric tag에 넣지 않는다.
|
|
25. regular request path에서 `KEYS`, unbounded `SCAN`, unbounded collection read를 실행하지 않는다.
|
|
26. unused capability는 connection, thread, scheduler, health dependency를 만들지 않는다.
|
|
27. `@Primary`, bean-name 존재만으로 provider와 guarantee를 선택하지 않는다.
|
|
28. provider cutover 중 JDBC와 Redis가 동시에 같은 scope를 독립 claim하게 하지 않는다.
|
|
29. liveness를 Redis availability에 연결하지 않는다.
|
|
30. real Redis/failure test 없이 R2/R3를 주장하지 않는다.
|
|
|
|
## 6. 대안 검토
|
|
|
|
### A. 현재 `RedisClient` seam에 method만 계속 추가
|
|
|
|
장점은 change surface가 작다는 것이다. 그러나 host SDK를 다시 추상화하는 거대한 low-level
|
|
interface가 되고 Redis semantics를 fake test로 흉내 내게 된다. Cluster redirect, Lua result,
|
|
timeout certainty, connection lifecycle을 새 interface가 부정확하게 복제한다.
|
|
|
|
선택하지 않는다.
|
|
|
|
### B. application에 범용 `RedisPort` 제공
|
|
|
|
예를 들어 `get/set/incr/zadd/eval`을 application port로 노출하면 개발자는 빠르게 기능을 만들 수
|
|
있다. 대신 use case가 provider key, TTL, serialization, data structure, atomic recipe를 직접
|
|
소유하고 Clean Architecture 경계가 무너진다. 안전 helper가 아닌 raw infrastructure facade가 된다.
|
|
|
|
선택하지 않는다.
|
|
|
|
### C. capability-provider마다 즉시 leaf 분리
|
|
|
|
```text
|
|
cache-redis
|
|
rate-limit-redis
|
|
lock-redis
|
|
idempotency-redis
|
|
session-redis
|
|
```
|
|
|
|
물리 격리는 가장 명확하다. 그러나 현재 exact-19 registry를 즉시 바꾸고 동일 client/config/program
|
|
기반을 여러 module에 중복한다. semantic contract가 아직 구현으로 검증되지 않은 시점에 public
|
|
path를 고정하는 비용이 크다.
|
|
|
|
R2 첫 단계에는 선택하지 않는다. 독립 release/security/dependency lifecycle이 생기면 다시
|
|
평가한다.
|
|
|
|
### D. 한 physical Redis leaf, capability별 package와 semantic port
|
|
|
|
현재 registry를 유지하면서 실제 SDK와 공통 key/codec/program runtime을 한 곳에 둘 수 있다.
|
|
동시에 capability별 provider, failure policy, settings, test kit를 분리할 수 있다.
|
|
|
|
선택한다. 단, “한 leaf”는 “한 connection”, “한 Redis deployment”, “한 fail-open policy”를
|
|
뜻하지 않는다.
|
|
|
|
### E. Redisson API를 중심으로 모든 기능 제공
|
|
|
|
Redisson은 lock, rate limiter, map cache 같은 고수준 primitive를 제공한다. 구현량은 줄지만
|
|
provider-specific semantics와 watchdog/failover 가정이 application policy에 스며들기 쉽고,
|
|
Spring Data/Spring Session과 별도 client lifecycle이 중복될 수 있다.
|
|
|
|
기본 선택으로 사용하지 않는다. 특정 product가 Redisson capability를 선택할 경우 동일 semantic
|
|
contract와 contract suite를 통과하는 별도 provider로 추가할 수 있다.
|
|
|
|
### F. Redis Functions만 허용
|
|
|
|
Functions는 server에 versioned library를 배포하고 runtime 계정에서 `FCALL`만 허용하기 쉬워
|
|
production least privilege에 유리하다. 그러나 일부 managed Redis의 provisioning 권한, version,
|
|
배포 lifecycle이 다르고 모든 primary에 선배포해야 한다.
|
|
|
|
production 권장 profile로 열어두지만 portable R2의 유일한 모드로 강제하지 않는다. EVALSHA
|
|
compatibility profile과 명시적으로 구분한다.
|
|
|
|
## 7. 목표 아키텍처
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
WEB[adapter:inbound:web] -->|HTTP mapping| EDGE[shared edge rate-limit contract]
|
|
WEB --> APP[application-core use case]
|
|
APP --> CACHEPORT[semantic cache region port]
|
|
APP --> IDEMPORT[idempotency port]
|
|
APP --> LEASEPORT[lease/fencing port]
|
|
|
|
CACHEPORT --> REDISCACHE[Redis cache provider]
|
|
IDEMPORT --> REDISIDEM[Redis idempotency provider]
|
|
IDEMPORT --> JPAIDEM[JPA idempotency provider]
|
|
LEASEPORT --> REDISLEASE[Redis lease provider]
|
|
LEASEPORT --> JDBCLOCK[JDBC lock provider]
|
|
EDGE --> REDISRATE[Redis rate-limit provider]
|
|
|
|
BOOT[app-bootstrap] -. selects/binds/validates .-> REDISCACHE
|
|
BOOT -. selects/binds/validates .-> REDISIDEM
|
|
BOOT -. selects/binds/validates .-> REDISLEASE
|
|
BOOT -. selects/binds/validates .-> REDISRATE
|
|
BOOT -. composes session mode .-> SESSION[Spring Session Redis]
|
|
|
|
REDISCACHE --> CACHEDEP[(cache deployment)]
|
|
REDISRATE --> COORDDEP[(coordination deployment)]
|
|
REDISIDEM --> COORDDEP
|
|
REDISLEASE --> COORDDEP
|
|
SESSION --> SESSIONDEP[(session deployment)]
|
|
```
|
|
|
|
핵심 방향은 다음과 같다.
|
|
|
|
```text
|
|
business/domain policy
|
|
-> semantic application port
|
|
-> Redis capability provider
|
|
-> internal key/codec/program/client runtime
|
|
-> role-bound Redis deployment
|
|
```
|
|
|
|
`adapter:inbound:web`와 `adapter:outbound:cache-redis`는 서로 직접 의존하지 않는다.
|
|
`shared-contract`가 transport-edge rate-limit value contract를 소유하고 `app-bootstrap`이 두
|
|
adapter를 조립한다.
|
|
|
|
## 8. 모듈과 계층 소유권
|
|
|
|
| 소유 leaf | 포함 | 금지 |
|
|
| --- | --- | --- |
|
|
| `domain-core` | 실제 domain invariant와 value | Redis/cache/session/HTTP quota |
|
|
| `application-core` | typed cache region base port, cache policy, idempotency v2, lease/fencing port | Redis key/SDK/Lua/Spring |
|
|
| `shared-contract` | edge rate-limit request/decision, provider-neutral capability descriptor | Servlet, Redis topology/key/program detail, business quota |
|
|
| `adapter:inbound:web` | route/principal/IP/policy 해석, HTTP header/error mapping, session security behavior | Redis command, local business policy |
|
|
| `adapter:outbound:cache-redis` | client, key, codec, program, cache/rate/idempotency/lease/session storage provider | controller/use case/business rule |
|
|
| `adapter:outbound:persistence-jpa` | JPA idempotency/JDBC lock provider | Redis provider fallback |
|
|
| `app-bootstrap` | provider selection, role binding, conditional composition, startup guarantee validation | use-case logic |
|
|
| `sample-portfolio` | 실제 사용 예와 contract fixture | production leaf의 역의존 |
|
|
|
|
현재 registry상 Redis leaf는 이미 `application-core`, `domain-core`, `shared-contract`,
|
|
`adapter-outbound-support`에 의존할 수 있다. 실제 구현 시 필요한 production edge만 Gradle에
|
|
추가하고 domain dependency가 불필요하면 추가하지 않는다.
|
|
|
|
### 8.1 Redis leaf package
|
|
|
|
초기 package 구조는 다음과 같다.
|
|
|
|
```text
|
|
dev.caskeleton.adapter.outbound.redis
|
|
runtime/
|
|
connection/
|
|
topology/
|
|
capability/
|
|
health/
|
|
key/
|
|
codec/
|
|
program/
|
|
cache/
|
|
ratelimit/
|
|
coordination/
|
|
lease/
|
|
fencing/
|
|
idempotency/
|
|
session/
|
|
observability/
|
|
config/
|
|
```
|
|
|
|
물리 path가 `cache-redis`여도 새 code의 package root는 기술 책임을 정직하게 드러내는
|
|
`...outbound.redis`를 사용한다. 기존 `...outbound.cache` package는 migration facade로 유지한 뒤
|
|
제거한다. package 변경은 public-path snapshot과 migration note를 동반한다.
|
|
|
|
### 8.2 leaf split trigger
|
|
|
|
다음 중 하나가 성립하면 별도 registry migration으로 split한다.
|
|
|
|
- Spring Session dependency를 cache/rate/lock consumer classpath에서 제거해야 한다.
|
|
- capability별 release cadence가 달라진다.
|
|
- 별도 security review와 artifact ownership이 필요하다.
|
|
- Redis Streams inbound consumer처럼 adapter direction이 바뀐다.
|
|
- client SDK가 달라진다.
|
|
- package-level ArchUnit만으로 dependency leakage를 막기 어렵다.
|
|
- module build/test 시간이 독립 lifecycle을 방해한다.
|
|
|
|
단순 class 수 증가는 split 근거가 아니다.
|
|
|
|
## 9. Capability readiness와 descriptor
|
|
|
|
Redis라는 기술 전체에 하나의 readiness를 붙이지 않는다.
|
|
|
|
| Level | 의미 | Redis 예 |
|
|
| --- | --- | --- |
|
|
| R0 | contract/seam | 현재 `RedisClient` |
|
|
| R1 | local service | standalone cache, local-only evidence |
|
|
| R2 | production baseline | real provider, security/failure/health/real-service test |
|
|
| R3 | scale/HA proven | failover/Cluster/rolling/capacity evidence |
|
|
|
|
descriptor는 두 층으로 나눈다. `shared-contract`는 bootstrap이 모든 provider에 공통으로 사용하는
|
|
provider-neutral descriptor만 소유한다.
|
|
|
|
```java
|
|
public record CapabilityDescriptor(
|
|
String capabilityId,
|
|
String providerId,
|
|
CapabilityReadiness readiness,
|
|
Set<Guarantee> guarantees,
|
|
Set<NonGuarantee> nonGuarantees,
|
|
FailureMode failureMode,
|
|
ReadinessImpact readinessImpact,
|
|
boolean multiInstanceCapable,
|
|
String implementationVersion) {}
|
|
```
|
|
|
|
Redis leaf는 provider-specific 진단을 별도 타입으로 소유한다.
|
|
|
|
```java
|
|
public record RedisProviderDescriptor(
|
|
String capabilityId,
|
|
RedisRole role,
|
|
RedisTopology topology,
|
|
String deploymentId,
|
|
String keySchemaVersion,
|
|
String hashKeyVersion,
|
|
String codecSchemaVersion,
|
|
String programSetVersion,
|
|
String minimumRedisVersion) {}
|
|
```
|
|
|
|
bootstrap의 provider selection은 `CapabilityDescriptor`만 사용한다. Redis-specific startup
|
|
validation과 bounded health detail만 `RedisProviderDescriptor`를 소비한다. 두 descriptor 모두
|
|
SDK object나 credential/endpoint를 포함하지 않는다.
|
|
|
|
예시:
|
|
|
|
```text
|
|
capability=cache.worklog-summary
|
|
provider=redis
|
|
implementationVersion=redis-cache-v2
|
|
readiness=R2
|
|
role=cache
|
|
guarantees=[BOUNDED_TTL, EXPLICIT_DEGRADED_RESULT]
|
|
nonGuarantees=[READ_YOUR_WRITES, ATOMIC_DB_CACHE_WRITE]
|
|
failureMode=FAIL_OPEN_TO_SOURCE
|
|
```
|
|
|
|
다음과 같은 descriptor는 startup에서 거절한다.
|
|
|
|
```text
|
|
capability=session
|
|
role=cache
|
|
failureMode=FAIL_OPEN
|
|
```
|
|
|
|
## 10. Redis role과 deployment isolation
|
|
|
|
### 10.1 최소 role
|
|
|
|
| Role | 데이터 | 기본 eviction | durability/read | 기본 장애 의미 |
|
|
| --- | --- | --- | --- | --- |
|
|
| `cache` | 재생성 가능한 positive/negative/stale entry와 `CACHE_REFRESH_SOFT_LEASE` | `allkeys-lfu` 또는 검증된 `allkeys-lru` | persistence optional, replica stale read optional | policy별 source fallback |
|
|
| `coordination` | strict quota, idempotency, lease/fence | `noeviction` | primary read, declared persistence/HA | fail closed/indeterminate |
|
|
| `session` | 인증 session과 optional index | `noeviction` | primary read, HA/persistence | re-auth 또는 fail closed |
|
|
|
|
높은 rate-limit volume이나 Streams workload가 noisy-neighbor가 되면 다음 role을 추가로 분리한다.
|
|
|
|
```text
|
|
rate-limit
|
|
stream
|
|
```
|
|
|
|
이는 key prefix나 Redis database number가 아니라 별도 managed database/cluster/instance를 뜻한다.
|
|
|
|
### 10.2 왜 prefix와 DB number로 충분하지 않은가
|
|
|
|
`maxmemory-policy`, CPU, event loop, persistence fork, replication buffer, failover, connection
|
|
limit은 instance/deployment 단위다. cache key가 eviction을 유발하면 같은 deployment의 session과
|
|
idempotency key도 정책의 영향을 받는다.
|
|
|
|
Redis Cluster는 database 0만 사용한다. standalone에서 DB 1, DB 2로 나누어도 memory와 failure
|
|
domain은 같다. 따라서 logical database는 namespace일 뿐 guarantee isolation이 아니다.
|
|
|
|
`CACHE_REFRESH_SOFT_LEASE`는 cache miss load를 줄이는 용도이고 eviction/loss/duplicate owner를
|
|
허용한다. 일반 `EFFICIENCY_LEASE`, idempotency, fencing, strict quota는 coordination role만
|
|
사용한다. 이 한정된 soft lease 예외를 generic lock binding으로 확대하지 않는다.
|
|
|
|
### 10.3 binding model
|
|
|
|
capability는 deployment endpoint를 직접 알지 못하고 role binding을 사용한다.
|
|
|
|
```text
|
|
cache region -> cache role -> cache-main deployment
|
|
strict rate -> coordination role -> coord-main deployment
|
|
idempotency -> coordination role -> coord-main deployment
|
|
session -> session role -> session-main deployment
|
|
```
|
|
|
|
role마다 connection factory와 client resources를 분리한다. 하나의 global
|
|
`RedisConnectionFactory @Primary`를 사용하지 않는다.
|
|
|
|
### 10.4 incompatible co-location validation
|
|
|
|
동일 physical deployment ID에 다음 조합이 bind되면 production startup을 거절한다.
|
|
|
|
- evictable cache + session;
|
|
- evictable cache + idempotency;
|
|
- evictable cache + fenced coordination;
|
|
- replica-read cache + primary-only correctness capability;
|
|
- mutually incompatible persistence/eviction attestation.
|
|
|
|
local profile은 명시적 `allow-unsafe-colocation=true`로만 한 container를 공유할 수 있으며
|
|
readiness는 R1로 강등된다.
|
|
|
|
### 10.5 deployment policy ownership
|
|
|
|
application은 `CONFIG SET`을 실행하지 않는다. maxmemory, eviction, AOF/RDB, replica,
|
|
Sentinel/Cluster, TLS, backup은 IaC/managed service 정책이 소유한다.
|
|
|
|
runtime은 가능한 경우 read-only introspection으로 effective policy를 확인한다. managed service가
|
|
`CONFIG GET`을 막으면 signed/operator attestation과 external conformance job을 사용한다. 확인할 수
|
|
없다는 이유로 원하는 guarantee가 존재한다고 추정하지 않는다.
|
|
|
|
## 11. Key model
|
|
|
|
### 11.1 canonical shape
|
|
|
|
모든 key는 중앙 `RedisKeyBuilder`로만 만든다.
|
|
|
|
```text
|
|
ca:<app>:<env>:<capability>:<region>:hv<hashKeyVersion>:kv<keyVersion>:{<slotTag>}:<resourceDigest>:<kind>
|
|
```
|
|
|
|
예시:
|
|
|
|
```text
|
|
ca:worklog-api:prod:cache:worklog-summary:hv1:kv2:{a8f3}:6eab...:entry
|
|
ca:worklog-api:prod:rate:login:hv1:kv1:{31d0}:98bd...:bucket
|
|
ca:worklog-api:prod:idem:create-worklog:hv2:kv2:{bf91}:9aa1...:record
|
|
ca:worklog-api:prod:lease:daily-export:hv1:kv1:{04cf}:2d50...:owner
|
|
```
|
|
|
|
`app`, `env`, capability와 region은 validated bounded slug다. tenant, principal, token,
|
|
session ID, client idempotency key, resource path는 raw로 넣지 않는다.
|
|
|
|
### 11.2 digest
|
|
|
|
식별자 종류에 따라 다음을 선택한다.
|
|
|
|
- 이미 random opaque ID이고 노출 위험이 낮음: bounded SHA-256 digest;
|
|
- 사용자/tenant/email/IP처럼 dictionary attack 가능한 값: versioned HMAC-SHA-256;
|
|
- composite scope: length-prefixed canonical encoding 후 HMAC;
|
|
- rate-limit IP: trusted resolver가 normalized binary address를 만들고 HMAC.
|
|
|
|
단순 문자열 delimiter join은 ambiguity가 있으므로 금지한다.
|
|
|
|
```text
|
|
len(tenant) || tenant || len(principal) || principal || len(operation) || operation
|
|
```
|
|
|
|
key HMAC secret은 payload encryption key와 분리한다. canonical key의
|
|
`hv<hashKeyVersion>` segment가 HMAC key version을 고정한다. rotation은 bounded
|
|
dual-read/dual-delete 또는 cold-cutover 정책을 명시하고, old `hv` key가 TTL/maintenance로
|
|
drain된 뒤 ACL pattern을 제거한다.
|
|
|
|
#### HMAC material과 rotation
|
|
|
|
`RedisKeyDigestMaterialProvider`는 Redis leaf 소유 SPI이고 app-bootstrap이 generic secret
|
|
provider를 bridge한다.
|
|
|
|
```java
|
|
public interface RedisKeyDigestMaterialProvider {
|
|
RedisKeyDigestMaterialResolution resolve(
|
|
KeyDigestProfileId profile, HashKeyVersion version, SecretReference reference);
|
|
RotationSubscription subscribe(
|
|
KeyDigestProfileId profile, RedisKeyDigestRotationListener listener);
|
|
}
|
|
|
|
public record VersionedRedisKeyDigestMaterial(
|
|
KeyDigestAlgorithm algorithm,
|
|
HashKeyVersion version,
|
|
Instant expiresAt,
|
|
DestroyableSecret keyBytes) {}
|
|
```
|
|
|
|
HMAC profile은 algorithm, one write version, bounded readable versions, version별 secret reference,
|
|
rotation mode를 모두 가져야 한다. resolve는 credential SPI와 같이 unavailable/expired/
|
|
permission/invalid를 구분하고 secret byte를 log/metric/descriptor에 넣지 않는다. Redis leaf가
|
|
`adapter:outbound:identifier` sibling에 의존하지 않는다.
|
|
|
|
rotation mode:
|
|
|
|
- `dual-read-delete`: cache처럼 재생성 가능한 data만 허용. write는 새 `hv`, read는 newest-first
|
|
bounded probe, old hit는 metric 후 new key로 refresh 가능, invalidate는 모든 readable `hv`를
|
|
bounded delete한다. 두 version의 key/slot을 atomic하다고 표현하지 않는다.
|
|
- `cold-cutover`: idempotency, lease/fence, strict rate처럼 두 namespace의 동시 owner/state가
|
|
위험한 capability. mutation admission을 닫고 holder/lease/state TTL을 drain/reconcile한 뒤
|
|
write/read version을 한 번에 전환한다.
|
|
- rate policy가 무중단 rotation을 요구하면 old/new limiter를 모두 평가해 어느 하나 deny면
|
|
deny하는 별도 conservative overlap revision을 사용한다. token/counter를 두 key 사이 atomic
|
|
migration했다고 주장하지 않는다.
|
|
- `fixed`: random opaque ID의 unkeyed digest처럼 secret rotation 축이 없는 profile.
|
|
|
|
startup은 write version material 존재/미만료, readable version 최대 개수, algorithm 일치,
|
|
capability에 허용된 rotation mode, ACL prefix/version을 검증한다. old material을 제거하기 전
|
|
key TTL upper bound, maintenance scan evidence, active owner/session 없음 또는 explicit cold
|
|
cutover evidence가 필요하다.
|
|
|
|
profile 정의만으로 secret을 resolve하거나 watcher를 시작하지 않는다. active capability가
|
|
profile을 참조할 때만 해당 version material/subscription을 만든다.
|
|
|
|
### 11.3 hash tag
|
|
|
|
`{slotTag}`는 같은 atomic operation에 필요한 최소 key group만 co-locate한다.
|
|
|
|
- idempotency record와 its operation marker;
|
|
- sliding counter의 current/previous bucket;
|
|
- lease owner와 fencing counter;
|
|
- exact rate decision dedup record.
|
|
|
|
tenant 전체를 hash tag로 쓰면 한 tenant의 모든 traffic이 한 slot/hot shard로 몰리므로 금지한다.
|
|
slot tag는 resource/policy digest의 bounded prefix다.
|
|
|
|
### 11.4 version
|
|
|
|
세 version을 분리한다.
|
|
|
|
```text
|
|
key schema version
|
|
payload schema version
|
|
policy revision
|
|
```
|
|
|
|
key version은 physical layout/namespace를 바꾼다. payload version은 같은 key의 decode
|
|
compatibility를 바꾼다. policy revision은 rate/cache TTL 등 state interpretation을 바꾼다.
|
|
|
|
정책이 바뀌었는데 기존 counter/token state를 새 의미로 재사용하지 않는다. rate-limit key에는
|
|
policy revision을 포함한다.
|
|
|
|
### 11.5 bounds
|
|
|
|
key builder는 다음을 검증한다.
|
|
|
|
- 전체 UTF-8 byte length;
|
|
- 각 slug length와 allowed character;
|
|
- digest algorithm/version;
|
|
- hash tag 정확히 하나;
|
|
- `{`, `}`가 user input에서 유입되지 않음;
|
|
- capability별 kind allowlist.
|
|
|
|
invalid key input은 backend outage가 아니며 fail-open하지 않는다.
|
|
|
|
### 11.6 mass invalidation
|
|
|
|
regular request에서 pattern delete를 하지 않는다.
|
|
|
|
선택지는:
|
|
|
|
1. key schema/version bump;
|
|
2. region generation ID 교체;
|
|
3. known-key bounded batch invalidation;
|
|
4. operator maintenance의 rate-limited `SCAN` + `UNLINK`.
|
|
|
|
generation key가 evict되어도 `0`으로 되돌아가 old namespace를 부활시키면 안 된다. missing이면
|
|
새 random 128-bit generation을 `SET NX`로 초기화하고 loser는 winner 값을 읽는다. old entry는
|
|
orphan이지만 다시 visible해지지 않는다.
|
|
|
|
## 12. Payload와 serialization
|
|
|
|
### 12.1 raw value 원칙
|
|
|
|
Redis runtime의 기본 value type은 `byte[]`다. application object를 reflection으로 자동
|
|
serialize하지 않는다. JDK native serialization과 unrestricted polymorphic/default typing은
|
|
금지한다.
|
|
|
|
application semantic port는 typed value를 사용하지만 adapter binding은 명시적 codec을 등록한다.
|
|
`CacheCodec`은 `adapter:outbound:cache-redis`의 provider SPI다. application use case는 이 타입을
|
|
보거나 호출하지 않는다.
|
|
|
|
```java
|
|
public interface CacheCodec<T> {
|
|
String schemaId();
|
|
int writeVersion();
|
|
byte[] encode(T value);
|
|
DecodeResult<T> decode(int storedVersion, byte[] payload);
|
|
}
|
|
```
|
|
|
|
`CacheCodec`에는 Jackson, JSON node, Redis serializer 타입이 없다. 구체 codec과
|
|
application-value mapping은 Redis leaf의 product-specific binding class가 소유한다. 현재
|
|
exact-19 skeleton에는 별도 application-adapter leaf가 없으므로 존재하지 않는 “mapping module”을
|
|
가정하지 않는다.
|
|
|
|
### 12.2 cache envelope
|
|
|
|
baseline envelope는 다음 field를 갖는다.
|
|
|
|
```text
|
|
magic
|
|
envelopeVersion
|
|
codecId
|
|
payloadVersion
|
|
flags [negative, compressed]
|
|
sourceRevision?
|
|
writtenAtEpochMillis
|
|
softExpiresAtEpochMillis?
|
|
hardExpiresAtEpochMillis
|
|
payloadLength
|
|
payloadDigest
|
|
payload
|
|
```
|
|
|
|
Redis key TTL은 hard expiry 이후의 physical cleanup을 담당한다. envelope hard expiry는 client가
|
|
stale/expired를 판정하고 clock/TTL drift를 관측하는 방어선이다.
|
|
|
|
### 12.3 compatibility
|
|
|
|
- writer는 한 version만 쓴다.
|
|
- reader는 현재 N과 migration window의 N-1을 읽는다.
|
|
- N-1 read는 N으로 opportunistic rewrite할 수 있다.
|
|
- unknown future version은 `SCHEMA_MISMATCH`이며 miss와 별도 metric을 남긴다.
|
|
- decoder exception, invalid length, digest mismatch는 `CORRUPT`다.
|
|
- corrupt entry는 bounded owner-safe quarantine/evict 후 policy에 따라 source를 조회한다.
|
|
- programming bug를 Redis unavailable로 분류하지 않는다.
|
|
|
|
rolling deploy에서 old reader가 new payload를 읽을 수 없으면 writer 전환 전에 dual-readable
|
|
codec을 배포한다.
|
|
|
|
### 12.4 size와 compression
|
|
|
|
region마다 다음 bound가 필수다.
|
|
|
|
- maximum encoded bytes;
|
|
- maximum decoded bytes;
|
|
- maximum collection elements;
|
|
- maximum compression ratio;
|
|
- encode/decode deadline.
|
|
|
|
oversize는 cache write를 `REJECTED_TOO_LARGE`로 만들 수 있으나 source result 자체를 실패시키지
|
|
않는다. session/idempotency response oversize는 해당 capability 계약에 따라 fail closed한다.
|
|
|
|
compression은 threshold 이상에서만 opt-in한다. decompression bomb를 막기 위해 decoded size와
|
|
ratio를 먼저 제한한다. secret과 attacker-controlled value를 같은 compressed context에 섞지
|
|
않는다.
|
|
|
|
### 12.5 session/idempotency codec
|
|
|
|
cache codec을 session과 idempotency에 그대로 재사용하지 않는다.
|
|
|
|
- session: allowlisted security/session attribute schema와 rolling compatibility;
|
|
- idempotency: request fingerprint metadata와 bounded response codec;
|
|
- rate/lease: fixed primitive schema, arbitrary object serialization 없음.
|
|
|
|
## 13. Atomic program registry
|
|
|
|
### 13.1 목적
|
|
|
|
Redis command 하나는 원자적으로 실행되지만 다음 client flow는 원자적이지 않다.
|
|
|
|
```text
|
|
GET -> decide -> SET
|
|
INCR -> EXPIRE
|
|
GET owner -> DEL
|
|
GET owner -> PEXPIRE
|
|
find record -> claim
|
|
```
|
|
|
|
다른 client가 명령 사이에 끼어들 수 있고 첫 command 성공 뒤 connection이 끊길 수 있다.
|
|
`AtomicRedisOperations`는 자주 필요한 안전 recipe를 adapter 내부에 제공한다.
|
|
|
|
### 13.2 application에 노출하지 않는 catalog
|
|
|
|
baseline:
|
|
|
|
```text
|
|
set-if-absent-with-ttl
|
|
compare-and-delete
|
|
compare-and-expire
|
|
compare-and-set-with-ttl
|
|
increment-with-initial-ttl
|
|
region-generation-init
|
|
region-generation-bump
|
|
cache-refresh-claim
|
|
cache-refresh-release
|
|
rate-fixed-window
|
|
rate-sliding-counter
|
|
rate-token-bucket
|
|
idempotency-claim
|
|
idempotency-start
|
|
idempotency-renew
|
|
idempotency-complete
|
|
idempotency-fail
|
|
idempotency-release
|
|
idempotency-inspect
|
|
idempotency-reconcile-committed
|
|
idempotency-reopen-no-effect
|
|
lease-acquire
|
|
lease-inspect
|
|
lease-renew
|
|
lease-release
|
|
session-create
|
|
session-inspect
|
|
session-save-if-live
|
|
session-touch-if-live
|
|
session-tombstone-and-delete
|
|
session-rotate
|
|
```
|
|
|
|
advanced:
|
|
|
|
```text
|
|
rate-sliding-log
|
|
rate-gcra
|
|
fenced-counter-provision
|
|
fenced-lease-acquire
|
|
fenced-lease-inspect
|
|
fenced-lease-renew
|
|
fenced-lease-release
|
|
bounded-semaphore-acquire
|
|
bounded-semaphore-release
|
|
stream-publish-with-dedup
|
|
```
|
|
|
|
catalog가 존재한다고 모든 application에서 사용해야 하는 것은 아니다. 사용하지 않는 program은
|
|
connection이나 state를 만들지 않는다.
|
|
|
|
### 13.3 descriptor
|
|
|
|
각 program은 code와 함께 다음 descriptor를 가진다.
|
|
|
|
```java
|
|
public record RedisProgramDescriptor(
|
|
String name,
|
|
int semanticVersion,
|
|
String libraryName,
|
|
String registeredFunctionName,
|
|
String sourceSha256,
|
|
int keyCount,
|
|
ClusterSlotRule clusterSlotRule,
|
|
String argumentSchema,
|
|
String resultSchema,
|
|
ComplexityBound complexity,
|
|
StateGrowthBound stateGrowth,
|
|
String minimumRedisVersion,
|
|
RetrySafety retrySafety,
|
|
TimeoutCertainty timeoutCertainty,
|
|
String metricOperation) {}
|
|
```
|
|
|
|
필수 문서:
|
|
|
|
- unsafe multi-command recipe;
|
|
- program이 보장하는 atomicity scope;
|
|
- 보장하지 않는 replication/durability;
|
|
- Redis server clock/client clock 사용 여부;
|
|
- maximum keys/arguments/value bytes/iterations;
|
|
- wrong type와 malformed state 처리;
|
|
- first write 전 validation;
|
|
- Cluster same-slot rule;
|
|
- timeout 이후 reconciliation 방법;
|
|
- backward/forward result compatibility.
|
|
|
|
### 13.4 bounded execution
|
|
|
|
Lua/Function은 실행 중 Redis의 다른 작업을 막는다. 따라서:
|
|
|
|
- O(1) 또는 명시적 작은 N;
|
|
- unbounded loop 금지;
|
|
- `KEYS`, dynamic `SCAN`, large `SMEMBERS/HGETALL/ZRANGE` 금지;
|
|
- caller가 collection bound를 우회하지 못하도록 server-side 검증;
|
|
- 모든 key를 `KEYS[]`로 전달;
|
|
- programmatically generated key 접근 금지;
|
|
- first mutation 전에 type, count, TTL, numeric range 검증;
|
|
- integer millisecond와 bounded fixed-point 사용;
|
|
- runtime source concatenation 금지;
|
|
- slow-program threshold와 CI execution budget 설정.
|
|
|
|
Redis transaction에는 rollback이 없고 script도 write 후 runtime error가 나면 partial effect에 대한
|
|
주의가 필요하다. program은 가능한 모든 검증을 첫 write 전에 끝낸다.
|
|
|
|
### 13.5 deployment mode
|
|
|
|
두 mode를 지원하되 자동 fallback하지 않는다.
|
|
|
|
#### `functions-provisioned`
|
|
|
|
- Redis Function library를 별도 provisioning job이 모든 primary에 선배포;
|
|
- `libraryName`과 `registeredFunctionName` 모두 semantic major/version을 포함;
|
|
- 예: `ca_rate_v2` library와 `ca_rate_token_bucket_v2` registered function;
|
|
- library name/version/digest 검증;
|
|
- application runtime 계정은 `FCALL`만 허용;
|
|
- v1/v2는 서로 다른 library와 registered function name으로 동시에 존재;
|
|
- v2 unique name을 모든 primary에 load -> replica propagation/all-node digest 확인 ->
|
|
application이 v2 registered name으로 전환 -> v1 caller drain 확인 -> v1 library delete;
|
|
- failover/reshard 후 promoted/new primary의 library 확인;
|
|
- startup mismatch는 required capability를 fail closed.
|
|
|
|
production least-privilege 권장 mode다.
|
|
|
|
동일 registered name을 유지해야 하면 blue/green 동시 존재를 주장하지 않는다.
|
|
`FUNCTION LOAD REPLACE`는 whole-library atomic replacement이므로 mixed-node rollout과 old caller
|
|
compatibility를 별도 절차로 다룬다.
|
|
|
|
#### `evalsha-managed`
|
|
|
|
- parameterized classpath Lua source와 SHA를 build artifact에 고정;
|
|
- `EVALSHA` 사용;
|
|
- `NOSCRIPT`일 때 fixed source만 `SCRIPT LOAD` 후 한 번 재시도;
|
|
- Cluster의 모든 target primary에서 lazy/eager load;
|
|
- runtime account가 `SCRIPT LOAD` 권한을 갖는 security trade-off 기록;
|
|
- pipeline 안에서 `NOSCRIPT` recovery를 기대하지 않음;
|
|
- source/checksum mismatch면 startup 실패.
|
|
|
|
managed service가 Functions를 지원하지 않을 때 쓰는 compatibility mode다.
|
|
|
|
Spring Data Redis 기본 script executor의 `EVALSHA -> EVAL` fallback에 의존하지 않는다. exact
|
|
`SCRIPT LOAD -> EVALSHA` protocol과 source allowlist를 유지하기 위해
|
|
`VersionedRedisProgramExecutor`가 dedicated/native connection callback을 소유한다. command trace
|
|
integration test가 dynamic `EVAL`이나 다른 source가 전송되지 않음을 검증한다.
|
|
|
|
### 13.6 program result
|
|
|
|
program은 ambiguous `0/1/null` 대신 versioned numeric tuple을 반환한다.
|
|
|
|
```text
|
|
[resultSchemaVersion, statusCode, serverNowMillis, primaryValue, auxiliaryValue...]
|
|
```
|
|
|
|
adapter가 status를 typed outcome으로 변환한다. unknown status/schema는 programming/
|
|
compatibility error이며 fail-open miss가 아니다.
|
|
|
|
### 13.7 mutation certainty
|
|
|
|
모든 mutation outcome은 최소 다음을 구분한다.
|
|
|
|
```text
|
|
APPLIED
|
|
NOT_APPLIED
|
|
CONFLICT
|
|
REJECTED
|
|
UNAVAILABLE_BEFORE_SEND
|
|
OVERLOADED_BEFORE_SEND
|
|
INDETERMINATE
|
|
```
|
|
|
|
socket timeout은 server가 command를 실행하지 않았다는 증거가 아니다. request가 server에 도달하고
|
|
response만 유실될 수 있다. `INDETERMINATE`는 operation token으로 inspect/reconcile하거나
|
|
capability-specific safe behavior로 전환한다.
|
|
|
|
### 13.8 normative `program-set.json`
|
|
|
|
program 이름 목록과 prose만으로 구현 호환성을 주장하지 않는다. 다음 machine-readable manifest가
|
|
program contract의 SSOT다.
|
|
|
|
```json
|
|
{
|
|
"programSet": "ca-redis-programs-v1",
|
|
"minimumRedisVersion": "7.2",
|
|
"resultSchemaVersion": 1,
|
|
"programs": [
|
|
{
|
|
"id": "compare-and-delete-v1",
|
|
"libraryName": "ca_primitive_v1",
|
|
"registeredFunctionName": "ca_compare_and_delete_v1",
|
|
"scriptResource": "redis/scripts/compare-and-delete-v1.lua",
|
|
"sha256": "<build-generated>",
|
|
"keys": [
|
|
{"index": 1, "name": "ownerKey", "sameSlotGroup": "resource"}
|
|
],
|
|
"arguments": [
|
|
{"index": 1, "name": "expectedOwner", "type": "opaque-bytes", "maxBytes": 128}
|
|
],
|
|
"state": {"type": "string", "maximumBytes": 128, "ttl": "existing"},
|
|
"validateBeforeFirstWrite": ["key-type", "owner-length"],
|
|
"statuses": ["DELETED", "ABSENT", "NOT_OWNER", "WRONG_TYPE"],
|
|
"complexity": "O(1)",
|
|
"stateGrowth": "none",
|
|
"clock": "none",
|
|
"retrySafety": "inspect-or-repeat-desired-absent",
|
|
"timeoutCertainty": "indeterminate",
|
|
"aclCommands": ["GET", "DEL"]
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
manifest schema 자체를 JSON Schema와 Java parser test로 고정한다. `<build-generated>`는 설계
|
|
placeholder가 아니라 build가 resource bytes에서 생성하고 release artifact에서 non-empty exact
|
|
digest로 치환해야 하는 field다.
|
|
|
|
각 entry는 반드시:
|
|
|
|
- exact `KEYS[index]`;
|
|
- exact ordered `ARGV[index]`, type, byte/numeric bound;
|
|
- state data type/field/version/TTL;
|
|
- first-write 이전 validation;
|
|
- complete status code enum;
|
|
- complexity와 per-call removal/iteration bound;
|
|
- clock source;
|
|
- retry/timeout certainty;
|
|
- minimum Redis version;
|
|
- exact ACL command allowlist;
|
|
- golden input/state/output vectors
|
|
|
|
를 갖는다. manifest, Java typed facade, Lua/Function source 중 하나라도 drift하면 build가 실패한다.
|
|
|
|
### 13.9 baseline helper normative matrix
|
|
|
|
아래는 baseline manifest의 필수 최소 shape다. `K1`, `A1`은 exact positional index다.
|
|
|
|
| Program | `KEYS` | Ordered `ARGV` | State / TTL | First-write boundary | Status |
|
|
| --- | --- | --- | --- | --- | --- |
|
|
| `set-if-absent-with-ttl-v1` | K1 target | A1 value bytes, A2 TTL ms, A3 operation ID | string / A2 | type, value bytes, TTL, op ID | `SET`, `EXISTS`, `WRONG_TYPE`, `INVALID` |
|
|
| `compare-and-delete-v1` | K1 owner | A1 expected owner | string / existing | type, owner bytes | `DELETED`, `ABSENT`, `NOT_OWNER`, `WRONG_TYPE` |
|
|
| `compare-and-expire-v1` | K1 owner | A1 expected owner, A2 new TTL ms | string / A2 on match | type, owner, TTL | `RENEWED`, `ABSENT`, `NOT_OWNER`, `WRONG_TYPE` |
|
|
| `compare-and-set-with-ttl-v1` | K1 entry | A1 expected revision/digest, A2 new envelope, A3 TTL ms | versioned string / A3 | type, envelope size/version, TTL | `STORED`, `ABSENT`, `REVISION_CONFLICT`, `INVALID` |
|
|
| `increment-with-initial-ttl-v1` | K1 counter | A1 positive delta, A2 max, A3 TTL ms | signed integer string / initialize once | type, range, overflow, TTL | `INCREMENTED`, `LIMIT_EXCEEDED`, `OVERFLOW`, `WRONG_TYPE` |
|
|
| `region-generation-init-v1` | K1 generation | A1 random generation ID | opaque string / no TTL | type, ID length | `INITIALIZED`, `EXISTING`, `WRONG_TYPE` |
|
|
| `region-generation-bump-v1` | K1 generation | A1 new random generation, A2 operation ID | opaque string / no TTL | type, generation/op length | `BUMPED`, `ALREADY_APPLIED`, `WRONG_TYPE` |
|
|
| `cache-refresh-claim-v1` | K1 soft lease | A1 owner, A2 lease TTL ms, A3 operation ID | owner envelope / A2 | type, owner/op, TTL | `ACQUIRED`, `CONTENDED`, `ALREADY_OWNED`, `INVALID` |
|
|
| `cache-refresh-release-v1` | K1 soft lease | A1 owner | owner envelope / existing | type, owner | `RELEASED`, `ABSENT`, `NOT_OWNER` |
|
|
|
|
공통 bound:
|
|
|
|
- key count는 표와 정확히 일치;
|
|
- value/envelope maximum은 region/provider manifest가 numeric value로 resolve;
|
|
- TTL `1..maxTtlMillis`;
|
|
- operation/owner ID는 fixed maximum bytes;
|
|
- `increment`는 signed 64-bit hard bound와 configured lower hard-fail threshold;
|
|
- generation/coordination non-ephemeral key는 role policy가 허용할 때만 TTL 없음.
|
|
|
|
### 13.10 capability state programs
|
|
|
|
| Program | `KEYS` | Ordered `ARGV` | State / TTL | First-write boundary | Status |
|
|
| --- | --- | --- | --- | --- | --- |
|
|
| `rate-fixed-window-v1` | K1 policy-subject | A1 policy revision, A2 limit, A3 cost, A4 window ms, A5 evaluation ID? | versioned hash / window remainder + grace | type/schema, ranges, clock regression, dedup | `ALLOWED`, `DENIED`, `DEDUP_REPLAY`, `CLOCK_UNSAFE`, `INVALID` |
|
|
| `rate-sliding-counter-v1` | K1 policy-subject | A1 revision, A2 limit, A3 cost, A4 window ms, A5 evaluation ID? | versioned hash / 2 windows + grace | same | same |
|
|
| `rate-token-bucket-v1` | K1 policy-subject | A1 revision, A2 capacity scaled, A3 refill scaled, A4 period ms, A5 cost scaled, A6 evaluation ID? | versioned hash / full-refill horizon + grace | type/schema, scale/ranges/overflow, clock | same |
|
|
| `idempotency-claim-v1` | K1 record | A1 fingerprint, A2 owner, A3 operation ID, A4 processing TTL ms, A5 replay TTL ms, A6 codec, A7 policy revision | versioned hash / state-dependent | complete request/state/type/size/TTL before owner write | `ACQUIRED`, `REPLAYED_ACQUIRE`, `COMPLETED_REPLAY`, `IN_PROGRESS`, `RECOVERY_REQUIRED`, `FINGERPRINT_MISMATCH`, `TAKEN_OVER_CLAIMED`, `OWNER_OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-start-v1` | K1 record | A1 owner, A2 attempt, A3 operation ID | `CLAIMED -> EXECUTING` / existing processing TTL | schema/state/owner/attempt/op before state write | `STARTED`, `ALREADY_STARTED_SAME_OPERATION`, `ABSENT`, `NOT_OWNER`, `NOT_CLAIMED`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-renew-v1` | K1 record | A1 owner, A2 attempt, A3 operation ID, A4 processing TTL ms | claimed/executing hash / renewed processing TTL | state/schema/owner/attempt/op/TTL | `RENEWED`, `ALREADY_RENEWED_SAME_OPERATION`, `ABSENT`, `NOT_OWNER`, `NOT_IN_PROGRESS`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-complete-v1` | K1 record | A1 owner, A2 attempt, A3 operation ID, A4 codec/version, A5 response digest, A6 response bytes/ref, A7 replay TTL ms | completed hash / replay TTL | all response/state/owner/attempt/op bounds before state write | `COMPLETED`, `ALREADY_COMPLETED_SAME_RESULT`, `RESPONSE_CONFLICT`, `ABSENT`, `NOT_OWNER`, `NOT_IN_PROGRESS`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-fail-v1` | K1 record | A1 owner, A2 attempt, A3 operation ID, A4 disposition, A5 retry/audit TTL ms | failed/abandoned hash / A5 | state/owner/attempt/op/disposition/TTL | `MARKED_RETRYABLE`, `MARKED_ABANDONED`, `ALREADY_MARKED_SAME_OPERATION`, `ABSENT`, `NOT_OWNER`, `NOT_IN_PROGRESS`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-release-v1` | K1 record | A1 owner, A2 attempt, A3 operation ID | claimed hash / delete or bounded audit marker | schema/state/owner/attempt/op before delete | `RELEASED_BEFORE_EXECUTION`, `ALREADY_RELEASED_SAME_OPERATION`, `ABSENT`, `NOT_OWNER`, `EXECUTION_ALREADY_STARTED`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-inspect-v1` | K1 record | A1 fingerprint, A2 owner, A3 operation ID | versioned hash / read-only with PTTL | schema/fingerprint/owner/op/state | `ABSENT`, `CLAIMED_SAME_OPERATION`, `EXECUTING_SAME_OPERATION`, `COMPLETED_REPLAY`, `IN_PROGRESS_OTHER`, `FAILED_RETRYABLE`, `ABANDONED`, `FINGERPRINT_MISMATCH`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `idempotency-reconcile-committed-v1` | K1 record | A1 expected attempt, A2 expected state revision, A3 evidence digest, A4 audit operation ID, A5 codec/version, A6 response digest, A7 response bytes/ref, A8 replay TTL ms | abandoned -> completed / replay TTL | schema/state/attempt/revision/evidence/response/TTL before state write | `RECONCILED_COMPLETED`, `ALREADY_RECONCILED_SAME_OPERATION`, `EVIDENCE_CONFLICT`, `STATE_CONFLICT`, `ABSENT`, `INVALID` |
|
|
| `idempotency-reopen-no-effect-v1` | K1 record | A1 expected attempt, A2 expected state revision, A3 evidence digest, A4 audit operation ID, A5 new owner, A6 new claim operation ID, A7 processing TTL ms | abandoned -> claimed attempt+1 / processing TTL | schema/state/attempt/revision/evidence/new owner/op/TTL before state write | `REOPENED_CLAIMED`, `ALREADY_REOPENED_SAME_OPERATION`, `EVIDENCE_CONFLICT`, `STATE_CONFLICT`, `ABSENT`, `INVALID` |
|
|
| `lease-acquire-v1` | K1 owner | A1 owner, A2 lease TTL ms, A3 operation ID | owner/op envelope / A2 | type/owner/op/TTL | `ACQUIRED`, `REPLAYED_SAME_OPERATION`, `CONTENDED`, `OWNER_OPERATION_CONFLICT`, `INVALID` |
|
|
| `lease-inspect-v1` | K1 owner | A1 owner, A2 operation ID | owner/op envelope / read-only with PTTL | type/owner/op/live TTL | `OWNED`, `ABSENT`, `NOT_OWNER`, `OWNER_OPERATION_CONFLICT`, `INVALID` |
|
|
| `lease-renew-v1` | K1 owner | A1 owner, A2 new TTL ms | owner envelope / A2 | type/owner/TTL | `RENEWED`, `ABSENT`, `NOT_OWNER`, `INVALID` |
|
|
| `lease-release-v1` | K1 owner | A1 owner | owner envelope / existing | type/owner | `RELEASED`, `ABSENT`, `NOT_OWNER` |
|
|
| `session-create-v1` | K1 session, K2 tombstone (same slot) | A1 mutation operation ID, A2 revision, A3 serializer/version, A4 payload digest, A5 payload, A6 idle TTL ms, A7 absolute expiry epoch ms | versioned session with last-mutation op/digest + tombstone / `min(idle, absolute-now)` | both types, op/payload/schema/revision/TTL/absolute before create | `CREATED`, `ALREADY_CREATED_SAME_OPERATION`, `EXISTS_CONFLICT`, `TOMBSTONED`, `ABSOLUTE_EXPIRED`, `INVALID` |
|
|
| `session-inspect-v1` | K1 session, K2 tombstone (same slot) | A1 expected mutation operation ID, A2 expected payload digest?, A3 expected revision? | live session/tombstone / read-only with PTTL | both schemas, last mutation/revision/digest/tombstone | `LIVE_SAME_MUTATION`, `LIVE_OTHER`, `TOMBSTONED_SAME_OPERATION`, `TOMBSTONED_OTHER`, `ABSENT`, `ABSOLUTE_EXPIRED`, `INVALID` |
|
|
| `session-save-if-live-v1` | K1 session, K2 tombstone (same slot) | A1 mutation operation ID, A2 expected revision, A3 new revision, A4 serializer/version, A5 payload digest, A6 payload, A7 idle TTL ms, A8 absolute expiry epoch ms | versioned session with last-mutation op/digest / bounded live TTL | both types, op/tombstone/revisions/payload/TTL/absolute before write | `SAVED`, `ALREADY_SAVED_SAME_OPERATION`, `ABSENT`, `STALE_REVISION`, `MUTATION_CONFLICT`, `TOMBSTONED`, `ABSOLUTE_EXPIRED`, `INVALID` |
|
|
| `session-touch-if-live-v1` | K1 session, K2 tombstone (same slot) | A1 mutation operation ID, A2 expected revision, A3 idle TTL ms, A4 absolute expiry epoch ms, A5 minimum touch interval ms | session last-touch op + bounded live TTL | both types, op/tombstone/revision/times before TTL/write | `TOUCHED`, `ALREADY_TOUCHED_SAME_OPERATION`, `TOUCH_NOT_DUE`, `ABSENT`, `STALE_REVISION`, `MUTATION_CONFLICT`, `TOMBSTONED`, `ABSOLUTE_EXPIRED`, `INVALID` |
|
|
| `session-tombstone-and-delete-v1` | K1 session, K2 tombstone (same slot) | A1 expected revision, A2 operation ID, A3 tombstone TTL ms | tombstone revision/op + deleted session / A3 | both types, revisions/op/TTL before tombstone first write | `REVOKED_AND_DELETED`, `TOMBSTONED_ABSENT`, `ALREADY_REVOKED_SAME_OPERATION`, `STALE_REVISION`, `OPERATION_CONFLICT`, `INVALID` |
|
|
| `session-rotate-v1` | K1 old session, K2 old tombstone, K3 new session, K4 new tombstone (same slot or non-Cluster profile) | A1 expected old revision, A2 new revision, A3 operation ID, A4 idle TTL ms, A5 absolute expiry epoch ms, A6 old tombstone TTL ms | new live session + old tombstone/delete / bounded TTLs | all four types, revisions/op/new-key absence/TTL/absolute before first write | `ROTATED`, `ALREADY_ROTATED_SAME_OPERATION`, `OLD_ABSENT`, `STALE_REVISION`, `OLD_TOMBSTONED`, `NEW_ID_CONFLICT`, `ABSOLUTE_EXPIRED`, `INVALID` |
|
|
| `fenced-counter-provision-v1` | K1 fence | A1 resource epoch, A2 registration digest, A3 durable high watermark | no-TTL versioned fence envelope | schema, epoch/digest, high watermark/range before create | `PROVISIONED`, `ALREADY_SAME`, `REGISTRATION_CONFLICT`, `REGRESSION`, `INVALID` |
|
|
| `fenced-lease-acquire-v1` | K1 owner, K2 fence (same slot) | A1 expected resource epoch, A2 registration digest, A3 owner, A4 lease TTL ms, A5 operation ID, A6 hard-fail threshold | owner/op/epoch/counter envelope + no-TTL versioned fence envelope | both types/schema/expected epoch+digest, owner/op, range, regression/missing policy, TTL before increment | `ACQUIRED`, `REPLAYED_SAME_OPERATION`, `CONTENDED`, `OWNER_OPERATION_CONFLICT`, `FENCE_COUNTER_MISSING`, `EPOCH_MISMATCH`, `REGISTRATION_CONFLICT`, `FENCE_REGRESSION`, `FENCE_EXHAUSTED`, `INVALID` |
|
|
| `fenced-lease-inspect-v1` | K1 owner, K2 fence (same slot) | A1 expected resource epoch, A2 registration digest, A3 owner, A4 operation ID | owner/op/epoch/counter envelope + no-TTL versioned fence envelope / read-only | both schemas/expected epoch+digest, owner/op, stored fence <= counter, live PTTL | `OWNED`, `ABSENT`, `NOT_OWNER`, `OWNER_OPERATION_CONFLICT`, `FENCE_COUNTER_MISSING`, `EPOCH_MISMATCH`, `REGISTRATION_CONFLICT`, `FENCE_REGRESSION`, `INVALID` |
|
|
| `fenced-lease-renew-v1` | K1 owner, K2 fence (same slot) | A1 expected epoch, A2 registration digest, A3 owner, A4 expected counter, A5 new TTL ms | same fenced owner envelope / A5; fence envelope no TTL | both schemas/epoch/digest/owner/counter/TTL before expire | `RENEWED`, `ABSENT`, `NOT_OWNER`, `TOKEN_MISMATCH`, `EPOCH_MISMATCH`, `REGISTRATION_CONFLICT`, `FENCE_COUNTER_MISSING`, `FENCE_REGRESSION`, `INVALID` |
|
|
| `fenced-lease-release-v1` | K1 owner, K2 fence (same slot) | A1 expected epoch, A2 registration digest, A3 owner, A4 expected counter | deleted owner; fence envelope unchanged/no TTL | both schemas/epoch/digest/owner/counter before delete | `RELEASED`, `ABSENT`, `NOT_OWNER`, `TOKEN_MISMATCH`, `EPOCH_MISMATCH`, `REGISTRATION_CONFLICT`, `FENCE_COUNTER_MISSING`, `FENCE_REGRESSION`, `INVALID` |
|
|
|
|
rate의 exact formula/boundary는 §20, idempotency state/result는 §24, fencing lifecycle은 §23이
|
|
추가 normative source다. manifest는 그 section의 revision/digest를 참조한다.
|
|
|
|
session program은 Redis `TIME`을 한 번 읽고 physical TTL을
|
|
`min(idleTtl, absoluteExpiresAt-serverNow)`로 정한다. tombstone/revision check와 save/touch/delete는
|
|
같은 invocation에서 이루어진다. Sentinel/standalone R2에서는 rotate 네 key를 한 primary에서
|
|
원자 실행한다. Cluster session profile은 네 key가 stable opaque session-lineage hash tag로 같은
|
|
slot임을 key builder/cookie format/`CLUSTER KEYSLOT` test가 증명할 때만 rotate guarantee를
|
|
광고한다. 그렇지 않으면 Cluster session profile은 R2가 아니다.
|
|
|
|
### 13.11 typed Java facade와 golden vectors
|
|
|
|
caller는 generic `execute(name, keys, args)`를 사용하지 않는다.
|
|
|
|
```java
|
|
CompareDeleteResult compareAndDelete(OwnerKey key, OwnerToken expected);
|
|
TokenBucketResult evaluateTokenBucket(TokenBucketCommand command);
|
|
IdempotencyCompleteResult complete(IdempotencyCompleteCommand command);
|
|
LeaseRenewResult renew(LeaseRenewCommand command);
|
|
```
|
|
|
|
각 facade의 sealed result enum은 manifest status와 1:1이다. unknown numeric status는
|
|
compatibility failure다.
|
|
|
|
각 program은 최소:
|
|
|
|
- empty/absent state;
|
|
- normal apply;
|
|
- condition reject;
|
|
- wrong owner/fingerprint/revision;
|
|
- boundary numeric/TTL;
|
|
- wrong Redis type;
|
|
- repeat same operation;
|
|
- timeout-after-apply reconciliation;
|
|
- Cluster same-slot
|
|
|
|
golden vector를 JSON fixture로 가진다. Functions와 EVALSHA mode가 같은 vector 결과를 내야 한다.
|
|
|
|
## 14. Command, transaction, pipeline 의미
|
|
|
|
### 14.1 single-thread 오해
|
|
|
|
Redis가 명령을 직렬 실행하더라도 client-side workflow 전체가 직렬화되는 것은 아니다.
|
|
|
|
```text
|
|
Client A: GET x
|
|
Client B: GET x
|
|
Client A: SET x 1
|
|
Client B: SET x 1
|
|
```
|
|
|
|
두 client 모두 동일 old value에서 판단할 수 있다. 단일 command, `WATCH` CAS, Lua/Function 중
|
|
하나로 atomic boundary를 만들어야 한다.
|
|
|
|
### 14.2 `MULTI/EXEC`
|
|
|
|
- queued command를 순서대로 실행한다.
|
|
- transaction 안의 다른 client command interleaving을 막는다.
|
|
- runtime command error에 rollback이 없다.
|
|
- network timeout 뒤 `EXEC` 수행 여부는 indeterminate일 수 있다.
|
|
- cross-slot key는 Cluster에서 사용할 수 없다.
|
|
|
|
`MULTI/EXEC`를 relational transaction으로 설명하지 않는다.
|
|
|
|
### 14.3 `WATCH`
|
|
|
|
`WATCH`는 optimistic CAS다. contention에서 abort/retry가 발생하고 connection affinity가 필요하다.
|
|
bounded low-contention update에는 사용할 수 있으나 rate limit, owner-safe release처럼 고빈도
|
|
primitive의 기본 구현은 Function/Lua를 사용한다.
|
|
|
|
### 14.4 pipeline
|
|
|
|
pipeline은 network round-trip을 줄인다. atomicity나 rollback을 제공하지 않는다. 다른 client의
|
|
command가 끼어들 수 있으며 partial response/timeout 처리도 필요하다.
|
|
|
|
다음 용도로 제한한다.
|
|
|
|
- independent bounded multi-get;
|
|
- independent bounded invalidation;
|
|
- metrics/maintenance의 bounded read;
|
|
- result별 성공/실패를 독립 처리할 수 있는 operation.
|
|
|
|
### 14.5 expirable mutation
|
|
|
|
다음은 금지한다.
|
|
|
|
```text
|
|
SET key value
|
|
EXPIRE key ttl
|
|
```
|
|
|
|
첫 command 후 process가 죽으면 TTL 없는 key가 남는다. `SET ... PX`, field TTL이 필요한
|
|
supported version command, 또는 atomic program을 사용한다.
|
|
|
|
### 14.6 common primitive 제공 원칙
|
|
|
|
개발자가 매번 raw command의 race와 bound를 다시 발견하게 두지 않는다. Redis leaf 내부에
|
|
`RedisPrimitiveCatalog`를 제공하되 application에 generic `RedisOperations`나 command string을
|
|
노출하지 않는다.
|
|
|
|
```text
|
|
StringValuePrimitives
|
|
CounterPrimitives
|
|
HashPrimitives
|
|
SetPrimitives
|
|
SortedSetPrimitives
|
|
ListPrimitives
|
|
BitmapPrimitives
|
|
HyperLogLogPrimitives
|
|
GeoPrimitives
|
|
```
|
|
|
|
각 primitive method는:
|
|
|
|
- typed/versioned key만 받음;
|
|
- encoded/decoded byte bound;
|
|
- finite deadline과 command certainty;
|
|
- role allowlist;
|
|
- Cluster slot rule;
|
|
- TTL policy;
|
|
- collection result/count bound;
|
|
- retry classification;
|
|
- metric operation name
|
|
|
|
을 descriptor에서 가져온다. raw `byte[] key`, arbitrary command, unbounded range, caller-provided
|
|
Lua source는 public API가 아니다.
|
|
|
|
새 product 기능이 이 primitive를 원하면 application에 `LeaderboardPort`,
|
|
`UniqueVisitorEstimatePort`처럼 semantic port를 만들고 Redis leaf의 thin adapter가 내부
|
|
primitive를 사용한다. core가 `ZADD`, `PFADD`를 직접 호출하지 않는다.
|
|
|
|
### 14.7 data structure별 baseline과 함정
|
|
|
|
| Structure | 제공 baseline | 반드시 막는 함정 |
|
|
| --- | --- | --- |
|
|
| String | bounded `GET`, `MGET`, `SET PX`, `SET NX/XX PX`, conditional delete/set | `SET` 후 별도 `EXPIRE`, unbounded value, GET-판단-SET race |
|
|
| Counter | bounded signed `INCRBY`, saturating/read, initial-TTL atomic program | `INCR` 후 별도 `EXPIRE`, overflow, response-loss automatic retry/double charge |
|
|
| Hash | `HGET`, bounded `HMGET`, `HSET`, `HDEL`, bounded `HSCAN` | request path `HGETALL`, unbounded fields, whole-key TTL을 field TTL로 오해 |
|
|
| Set | `SISMEMBER`, `SADD`, `SREM`, `SCARD`, bounded `SSCAN` | unbounded `SMEMBERS`, attacker-controlled cardinality, exact set algebra on huge keys |
|
|
| Sorted set | `ZADD`, `ZREM`, `ZCOUNT`, rank/score range with explicit limit, bounded trim | unbounded `ZRANGE`, floating score precision 오해, trim의 `O(log N + M)` 비용 |
|
|
| List | bounded push/pop/trim, blocking pop은 전용 connection | durable queue/ack/reclaim로 오해, unbounded `LRANGE`, shared connection block |
|
|
| Bitmap | bounded `GETBIT`/`SETBIT`/`BITCOUNT`, fixed offset domain | attacker가 큰 offset으로 sparse allocation 유발, tenant bit leakage |
|
|
| HyperLogLog | `PFADD`/`PFCOUNT`를 approximate cardinality semantic port 뒤에서 사용 | exact count/billing/security decision에 사용, uncontrolled merge fan-in |
|
|
| Geo | bounded `GEOADD`/radius search with count/sort bound | raw precise location 보존·로그, unbounded radius/result, authorization 누락 |
|
|
|
|
Redis 7.2 portable minimum에서는 hash field별 expiration을 baseline으로 가정하지 않는다. hash
|
|
field마다 독립 TTL이 필요하면:
|
|
|
|
1. field를 독립 versioned key로 분리하거나;
|
|
2. envelope expiry + bounded lazy cleanup program을 사용하거나;
|
|
3. product minimum version을 별도 ADR/evidence로 높인다.
|
|
|
|
새er command가 더 편리해도 compatibility matrix가 지원하기 전에는 program manifest의 portable
|
|
recipe를 유지한다.
|
|
|
|
### 14.8 multi-step recipe의 atomic 승격 기준
|
|
|
|
다음 형태는 single-thread Redis에서도 client race가 있으므로 제공 primitive가 Function/Lua 또는
|
|
검증된 `WATCH` CAS로 승격한다.
|
|
|
|
| Unsafe recipe | Safe primitive |
|
|
| --- | --- |
|
|
| `GET owner` -> compare -> `DEL` | `compare-and-delete-v1` |
|
|
| `GET owner` -> compare -> `PEXPIRE` | `compare-and-expire-v1` |
|
|
| `INCR` -> first request면 `EXPIRE` | `increment-with-initial-ttl-v1` |
|
|
| `GET version` -> compare -> `SET PX` | `compare-and-set-with-ttl-v1` |
|
|
| `SCARD` -> limit check -> `SADD` | bounded set-admission program |
|
|
| `ZREMRANGEBYSCORE` -> `ZCOUNT` -> `ZADD` | rate/sliding-log program |
|
|
| `LLEN` -> capacity check -> `LPUSH` | bounded list-admission program |
|
|
| `HGET revision` -> conditional `HSET` | revision-CAS program |
|
|
|
|
program을 쓴다고 무조건 안전한 것은 아니다. §13 manifest의 first-write validation, time/TTL,
|
|
same-slot, state-growth, result schema, timeout certainty를 모두 가져야 catalog에 등록된다.
|
|
|
|
### 14.9 bulk와 collection API
|
|
|
|
bulk operation은 `maximumKeys`, `maximumElements`, `maximumEncodedBytes`, total deadline을
|
|
필수로 받거나 region descriptor에서 고정한다.
|
|
|
|
- `MGET`/pipeline은 같은 snapshot이나 atomic read가 아니다.
|
|
- Cluster multi-key command는 same-slot일 때만 사용하고, 그 외에는 bounded per-node pipeline과
|
|
partial result를 반환한다.
|
|
- bulk mutation의 일부 성공을 단일 boolean로 합치지 않는다.
|
|
- `SCAN` 결과는 duplicate/missing observation을 허용하는 maintenance cursor다.
|
|
- collection page token은 topology/schema revision을 포함하고 unlimited export API가 아니다.
|
|
|
|
### 14.10 queue와 messaging 경계
|
|
|
|
List의 `LPUSH/BRPOP`만으로 ack, visibility timeout, reclaim, poison handling, durable replay가
|
|
생기지 않는다. 단순 best-effort work handoff가 아니라면:
|
|
|
|
- Redis Streams의 consumer group/PENDING/claim/ack/trim을 가진 별도 messaging semantic port;
|
|
- 또는 Kafka/JDBC queue/outbox
|
|
|
|
를 선택한다. Stream도 duplicate delivery, pending-entry leak, trim data loss, consumer crash,
|
|
Cluster key placement를 별도 contract로 해결해야 하며 cache primitive catalog가 durable
|
|
messaging guarantee를 광고하지 않는다.
|
|
|
|
### 14.11 Redis Stack/module capability
|
|
|
|
Bloom/Cuckoo filter, Count-Min Sketch, Top-K, TimeSeries, JSON, Search/vector query는 plain Redis 7.2
|
|
portable command가 아니다. skeleton baseline에 있는 것처럼 보이게 하지 않고 각각 opt-in
|
|
provider capability로 연다.
|
|
|
|
| Capability | Semantic contract 예 | 핵심 non-guarantee/risk |
|
|
| --- | --- | --- |
|
|
| Probabilistic membership | `MightContainPort` | false positive, capacity/error-rate sizing, rebuild |
|
|
| Approximate frequency | `FrequencyEstimatePort` | exact billing/audit 불가, merge/error bound |
|
|
| Time series | `MetricSeriesPort` | retention/downsample/duplicate policy, observability backend 대체 아님 |
|
|
| JSON document | product-specific document port | aggregate consistency/JPA replacement 아님, schema/index migration |
|
|
| Search/vector | `SearchPort`/retrieval port | eventual index visibility, ranking drift, memory/index rebuild |
|
|
|
|
activation은 exact module/server image, command/version compatibility, ACL, license, backup/restore,
|
|
Cluster/failover, memory amplification, index build/rolling migration test가 있는 capability card를
|
|
요구한다. module이 없는 server에서 command probe 실패 시 plain data structure로 자동
|
|
fallback하지 않는다.
|
|
|
|
## 15. Cache application contract
|
|
|
|
### 15.1 port 형태
|
|
|
|
application-core는 technical TTL/codec/topology를 모르는 최소 provider-neutral base contract를
|
|
제공한다.
|
|
|
|
```java
|
|
public interface CacheRegionPort<K, V> {
|
|
CacheLookup<V> lookup(K key);
|
|
CacheRecordOutcome record(K key, V value, CacheRecordMetadata metadata);
|
|
CacheRecordOutcome recordAbsent(
|
|
K key, AuthoritativeAbsence reason, CacheRecordMetadata metadata);
|
|
CacheInvalidationOutcome invalidate(K key);
|
|
}
|
|
```
|
|
|
|
실제 use case는 semantic name을 갖는 interface를 정의한다.
|
|
|
|
```java
|
|
public interface WorkLogSummaryCachePort
|
|
extends CacheRegionPort<WorkLogSummaryKey, WorkLogSummarySnapshot> {}
|
|
```
|
|
|
|
위 이름은 forked product의 illustrative shape이며 현재 production template이나
|
|
`sample-portfolio`에 추가하지 않는다. current registry에는
|
|
`sample-portfolio -> adapter-outbound-cache-redis` edge가 없고 추가하지 않는다.
|
|
|
|
실제 product에서는 `application-core`에 semantic subtype을 두고 Redis leaf가 thin binding을
|
|
명시적으로 구현한다.
|
|
|
|
```java
|
|
final class WorkLogSummaryRedisCacheAdapter implements WorkLogSummaryCachePort {
|
|
private final RedisCacheRegion<WorkLogSummaryKey, WorkLogSummarySnapshot> delegate;
|
|
// 모든 method를 delegate하되 region ID/codec/effective policy는 constructor에서 freeze한다.
|
|
}
|
|
```
|
|
|
|
Redis leaf는 `application-core` dependency가 registry에서 허용되어 이 방향이 합법적이다.
|
|
generic `RedisCacheRegion<K,V>` bean 하나가 subtype을 자동 구현한다고 가정하지 않는다. skeleton의
|
|
실행 가능한 reference는 sample dependency가 아니라 Redis leaf test source의 test-only semantic
|
|
port/binding/codec fixture로 증명한다.
|
|
|
|
### 15.2 lookup
|
|
|
|
```text
|
|
Hit(
|
|
value,
|
|
freshness = FRESH | STALE,
|
|
sourceRevision?,
|
|
softExpiresAt?,
|
|
hardExpiresAt
|
|
)
|
|
|
|
NegativeHit(
|
|
reason,
|
|
hardExpiresAt
|
|
)
|
|
|
|
Miss(
|
|
reason = ABSENT | EXPIRED | INVALIDATED
|
|
)
|
|
|
|
IncompatibleSchema(
|
|
category = FUTURE_VERSION | RETIRED_VERSION | UNKNOWN_ENVELOPE,
|
|
policy = FAIL_FAST | QUARANTINE_AND_RELOAD
|
|
)
|
|
|
|
Unavailable(
|
|
category = UNAVAILABLE | OVERLOADED,
|
|
certainty
|
|
)
|
|
```
|
|
|
|
topology, OOM, connection 같은 provider detail은 Redis adapter metric/log에 남고 application
|
|
contract에는 노출하지 않는다.
|
|
|
|
`SCHEMA_MISMATCH`, `CORRUPT`, `PROGRAMMING_ERROR`는 ordinary unavailable이나 miss로 반환하지
|
|
않는다. future writer version은 default `FAIL_FAST` + readiness/compatibility alert이며 old reader가
|
|
entry를 지우거나 source로 덮어쓰지 않는다. approved retired-version migration처럼 region의
|
|
compiled policy가 `QUARANTINE_AND_RELOAD`를 명시한 경우에만 bounded quarantine/invalidate 후
|
|
source load를 허용하고, 결과/metric은 계속 `IncompatibleSchema` 경로로 기록한다. corrupt와
|
|
programming error는 typed fatal result/exception과 alert를 사용한다.
|
|
|
|
### 15.3 record metadata와 policy SSOT
|
|
|
|
```java
|
|
public record CacheRecordMetadata(
|
|
String sourceRevision,
|
|
CacheRecordIntent intent) {}
|
|
```
|
|
|
|
`CacheRecordIntent`는 `UPSERT` 또는 `ONLY_IF_SOURCE_REVISION_NEWER`처럼 application-visible
|
|
consistency 의도만 표현한다.
|
|
|
|
TTL, jitter, maximum bytes, codec, compression, technical retry는 region descriptor가 유일한
|
|
SSOT다. caller가 invocation마다 override하지 않는다.
|
|
|
|
startup의 `CacheRegionPolicyCompiler`가:
|
|
|
|
```text
|
|
code-declared semantic policy
|
|
+ environment operational bounds
|
|
+ provider capability limits
|
|
-> immutable EffectiveCacheRegionPolicy(revision, digest)
|
|
```
|
|
|
|
를 만들고 binding adapter에 freeze한다. rolling request마다 policy가 바뀌지 않는다.
|
|
|
|
### 15.4 mutation outcome
|
|
|
|
```text
|
|
RECORDED
|
|
NOT_RECORDED_CONDITION
|
|
NOT_RECORDED_PROVIDER_POLICY
|
|
DEGRADED_UNAVAILABLE
|
|
INDETERMINATE
|
|
```
|
|
|
|
cache-aside load 성공 후 cache write가 unavailable이어도 source response는 보통 성공한다. 그러나
|
|
metric과 degraded result는 남긴다. oversize/TTL/codec 같은 technical reason은 adapter
|
|
telemetry에 있고 application은 provider policy상 record되지 않았다는 사실만 본다.
|
|
codec/programming error는 이 downgrade 대상이 아니다.
|
|
|
|
### 15.5 cache-aside executor
|
|
|
|
각 use case가 같은 cache recipe를 다시 구현하지 않도록 application-core에 framework-free
|
|
`CacheAsideExecutor`를 제공한다.
|
|
|
|
```java
|
|
CacheResult<V> getOrLoad(
|
|
K key,
|
|
CacheRegionPort<K, V> region,
|
|
CacheSourceLoader<K, V> sourceLoader)
|
|
```
|
|
|
|
region-specific compiled application policy는 executor construction 시 주입되고 호출마다 전달하지
|
|
않는다.
|
|
|
|
```java
|
|
public interface CacheSourceLoader<K, V> {
|
|
SourceLoadOutcome<V> load(K key, CancellationToken cancellation);
|
|
}
|
|
|
|
public sealed interface SourceLoadOutcome<V> {
|
|
record Loaded<V>(V value, String sourceRevision) implements SourceLoadOutcome<V> {}
|
|
record AuthoritativeAbsent<V>(
|
|
AuthoritativeAbsence reason, String sourceRevision) implements SourceLoadOutcome<V> {}
|
|
record TransientFailure<V>(SourceFailure failure) implements SourceLoadOutcome<V> {}
|
|
record PermanentFailure<V>(SourceFailure failure) implements SourceLoadOutcome<V> {}
|
|
record Cancelled<V>() implements SourceLoadOutcome<V> {}
|
|
}
|
|
```
|
|
|
|
`SourceFailure`은 bounded application error category/code와 original cause를 보존하되 cause
|
|
message를 Redis/log/tag에 serialize하지 않는다. unclassified thrown exception은
|
|
`PermanentFailure`처럼 조용히 cache하지 않고 원래 예외를 보존해 전파한다.
|
|
|
|
`CacheResult<V>`는 최소:
|
|
|
|
```text
|
|
FreshHit
|
|
StaleHit
|
|
LoadedFromSource
|
|
AuthoritativeAbsent
|
|
StaleFallbackAfterTransientFailure
|
|
DegradedSourceResult
|
|
```
|
|
|
|
를 구분한다.
|
|
|
|
executor가 소유한다.
|
|
|
|
- lookup outcome 해석;
|
|
- source fallback;
|
|
- local single-flight;
|
|
- `AuthoritativeAbsent`만 negative cache;
|
|
- transient/permanent source failure를 negative cache하지 않음;
|
|
- stale-if-error;
|
|
- refresh claim;
|
|
- write outcome 기록 hook;
|
|
- caller cancellation/deadline 전파.
|
|
|
|
executor가 소유하지 않는다.
|
|
|
|
- Redis serialization;
|
|
- database transaction;
|
|
- domain authorization;
|
|
- source error/absence classification의 business rule;
|
|
- HTTP response mapping.
|
|
|
|
### 15.6 region descriptor
|
|
|
|
각 region은 code/config의 typed descriptor로 선언한다.
|
|
|
|
```text
|
|
regionId
|
|
required/optional
|
|
value schema/codec
|
|
positive TTL
|
|
negative TTL
|
|
soft/hard TTL
|
|
jitter
|
|
maximum payload
|
|
failure mode
|
|
stale-if-error
|
|
stampede strategy
|
|
invalidation strategy
|
|
read consistency
|
|
metrics cardinality key
|
|
```
|
|
|
|
arbitrary runtime user input으로 region을 만들지 않는다. startup에서 binding과 codec uniqueness를
|
|
검증한다.
|
|
|
|
## 16. Cache strategy catalog
|
|
|
|
### 16.1 cache-aside baseline
|
|
|
|
```text
|
|
lookup
|
|
HIT -> return
|
|
MISS -> load source -> store -> return
|
|
UNAVAILABLE -> policy에 따라 source load -> degraded return
|
|
```
|
|
|
|
source of truth는 Redis가 아니다. source loader failure와 cache failure를 별도로 분류한다.
|
|
|
|
장점:
|
|
|
|
- 명확한 ownership;
|
|
- only-read 데이터에 적합;
|
|
- Redis outage에서 source fallback 가능.
|
|
|
|
위험:
|
|
|
|
- miss burst;
|
|
- stale entry;
|
|
- DB와 cache dual-write gap;
|
|
- source overload.
|
|
|
|
따라서 single-flight, TTL jitter, bounded fallback, invalidation을 함께 설계한다.
|
|
|
|
### 16.2 negative cache
|
|
|
|
다음처럼 “존재하지 않음”이 source에서 확정된 경우만 cache한다.
|
|
|
|
- authoritative not-found;
|
|
- deterministic empty query;
|
|
- permission과 무관한 public absence.
|
|
|
|
다음은 negative cache하지 않는다.
|
|
|
|
- timeout;
|
|
- 5xx;
|
|
- authorization denial을 다른 principal과 공유;
|
|
- transient replication lag;
|
|
- validation/programming error.
|
|
|
|
negative TTL은 positive TTL보다 짧고 별도 policy다. attacker가 random key로 negative entry를
|
|
폭증시키지 못하도록 subject normalization, admission policy, cardinality budget을 둔다.
|
|
|
|
### 16.3 stale-while-revalidate
|
|
|
|
entry는 soft/hard expiry를 갖는다.
|
|
|
|
```text
|
|
now < soft -> FRESH
|
|
soft <= now < hard -> STALE, 한 worker refresh
|
|
hard <= now -> MISS, source load 필요
|
|
```
|
|
|
|
stale data를 반환해도 되는 query에만 사용한다. authorization, balance, inventory reservation,
|
|
revocation처럼 stale가 위험한 데이터에는 적용하지 않는다.
|
|
|
|
### 16.4 stale-if-error
|
|
|
|
source load가 transient failure일 때 hard expiry 전 stale 값을 반환할 수 있다. 반환 결과에는
|
|
`source=STALE_FALLBACK`, age, policy revision을 내부적으로 남긴다.
|
|
|
|
stale 최대 age를 무한히 연장하지 않는다. Redis write 실패를 이유로 hard expiry를 client에서
|
|
임의 연장하지 않는다.
|
|
|
|
### 16.5 refresh-ahead
|
|
|
|
read traffic이 없어도 반드시 warm해야 하는 bounded hot set에만 사용한다.
|
|
|
|
- region이 refresh 대상 key 목록을 소유;
|
|
- scheduler queue bounded;
|
|
- per-key single-flight;
|
|
- shutdown cancellation;
|
|
- source failure backoff;
|
|
- full keyspace `SCAN`으로 대상 발견 금지.
|
|
|
|
일반 cache의 기본은 아니다.
|
|
|
|
### 16.6 probabilistic early refresh
|
|
|
|
hot key가 동시에 soft expiry에 도달하는 것을 줄이기 위해 remaining TTL, prior load duration,
|
|
bounded random 값을 사용해 일부 request만 일찍 refresh한다.
|
|
|
|
이 전략은 correctness가 아니라 load smoothing이다. 확률 식과 upper bound를 descriptor에
|
|
versioning하고 deterministic property test를 둔다.
|
|
|
|
### 16.7 write-through
|
|
|
|
application write와 cache write를 함께 호출할 수 있으나 DB와 Redis가 한 transaction이라는 뜻은
|
|
아니다.
|
|
|
|
```text
|
|
DB commit succeeds
|
|
Redis update fails
|
|
```
|
|
|
|
경로가 존재한다. baseline은 blind value update보다 after-commit invalidate를 선호한다.
|
|
|
|
### 16.8 write-behind
|
|
|
|
in-memory executor가 Redis/DB에 나중에 쓰는 형태는 production baseline이 아니다. write-behind가
|
|
필요하면 durable outbox/stream, retry, ordering, terminal failure, reconciliation을 가진 별도
|
|
workflow로 설계한다.
|
|
|
|
### 16.9 L1 local + L2 Redis
|
|
|
|
optional profile:
|
|
|
|
```text
|
|
request -> bounded local L1 -> Redis L2 -> source
|
|
```
|
|
|
|
요구사항:
|
|
|
|
- L1 maximum weight와 expiry;
|
|
- L1 entry는 L2 hard expiry를 넘지 않음;
|
|
- invalidation disconnect 시 L1 전체 flush;
|
|
- per-region enable;
|
|
- local stale age metric;
|
|
- Pub/Sub/tracking 유실 시 TTL recovery;
|
|
- session/idempotency/strict rate에 적용 금지.
|
|
|
|
### 16.10 admission
|
|
|
|
모든 source result를 cache하지 않는다.
|
|
|
|
- payload size;
|
|
- expected reuse;
|
|
- load cost;
|
|
- tenant fairness;
|
|
- error/negative classification;
|
|
- sensitive data classification;
|
|
- cardinality budget.
|
|
|
|
low-reuse high-cardinality scan 결과를 admission하지 않아 cache pollution을 줄인다.
|
|
|
|
## 17. Cache invalidation과 consistency
|
|
|
|
### 17.1 source of truth
|
|
|
|
Redis cache는 source of truth가 아니다. cache read consistency는 region descriptor에 다음 중
|
|
하나로 표시한다.
|
|
|
|
```text
|
|
BEST_EFFORT
|
|
BOUNDED_STALENESS
|
|
READ_AFTER_INVALIDATION
|
|
SOURCE_REVISION_GUARDED
|
|
```
|
|
|
|
`STRONG`이나 `LINEARIZABLE`은 일반 DB + Redis cache 조합에 제공하지 않는다.
|
|
|
|
### 17.2 write ordering
|
|
|
|
금지:
|
|
|
|
```text
|
|
cache delete
|
|
DB transaction
|
|
```
|
|
|
|
DB transaction이 rollback하면 유효 entry만 제거되어 불필요한 load가 생긴다. 더 위험한 구현은
|
|
transaction 안에서 cache를 update한 뒤 DB가 rollback하는 것이다.
|
|
|
|
baseline:
|
|
|
|
```text
|
|
DB transaction commit
|
|
after-commit cache invalidate
|
|
```
|
|
|
|
process가 commit 직후 죽으면 invalidate가 누락될 수 있으므로 TTL이 최종 복구 경계다.
|
|
|
|
### 17.3 reliable invalidation
|
|
|
|
bounded stale만으로 충분하지 않으면 DB transaction에 invalidation intent를 outbox로 함께 append한다.
|
|
|
|
```text
|
|
business update + outbox invalidation intent -- same DB transaction
|
|
outbox relay -> cache invalidation consumer
|
|
```
|
|
|
|
outbox/messaging leaf와 Redis leaf가 서로 의존하지 않는다. application event/intent와 bootstrap
|
|
composition을 통해 연결한다.
|
|
|
|
CDC도 사용할 수 있다. 그러나 table-change를 cache key로 변환하는 mapping, ordering,
|
|
checkpoint, replay, schema evolution을 별도 consumer가 소유한다. “Debezium을 붙이면 cache
|
|
consistency가 해결된다”고 설명하지 않는다.
|
|
|
|
### 17.4 stale refill race
|
|
|
|
다음 race를 고려한다.
|
|
|
|
```text
|
|
R1: cache miss
|
|
R1: old DB value load
|
|
W : DB update + cache invalidate
|
|
R1: old value cache put
|
|
```
|
|
|
|
단순 delete는 stale value를 다시 채울 수 있다. 해결 선택지는:
|
|
|
|
1. source revision/version을 entry에 저장하고 conditional put;
|
|
2. generation ID를 read 시작 시 capture하고 같은 generation에만 put;
|
|
3. invalidation event의 revision보다 오래된 put 거절;
|
|
4. 아주 짧은 TTL로 bounded risk 수용.
|
|
|
|
baseline R2는 source가 revision을 제공할 수 있으면 `REPLACE_IF_SOURCE_REVISION`, 그렇지 않으면
|
|
generation capture를 사용한다.
|
|
|
|
### 17.5 generation protocol
|
|
|
|
```text
|
|
g = read/current generation
|
|
lookup key(g, logicalKey)
|
|
load source
|
|
put key(g, logicalKey) // captured g, current generation 재조회 금지
|
|
```
|
|
|
|
invalidation은 generation을 새 random value로 바꾼다. load 중 invalidation이 일어나면 old
|
|
generation에 stale put이 되지만 new readers는 새 generation만 본다.
|
|
|
|
generation bump response가 유실되면 result는 `INDETERMINATE`다. caller는 current generation을
|
|
inspect하고 desired operation token의 audit record가 있으면 reconcile한다. cache correctness가
|
|
TTL로 충분한 region은 duplicate bump를 허용하되 cold-cache 영향만 기록할 수 있다.
|
|
|
|
### 17.6 per-key invalidate
|
|
|
|
per-key invalidation은 idempotent delete다. timeout 후 재시도해도 최종 absent가 목적이므로 retry
|
|
safety가 높다. 단, stale refill race를 막는 revision/generation protocol과 함께 써야 한다.
|
|
|
|
large value는 `DEL`의 synchronous deallocation latency를 피하기 위해 supported deployment에서
|
|
`UNLINK`를 사용할 수 있다. key 존재 여부를 authoritative receipt로 해석하지 않는다.
|
|
|
|
### 17.7 Pub/Sub invalidation
|
|
|
|
Pub/Sub과 keyspace notification은 best-effort hint다.
|
|
|
|
- subscriber disconnect 중 event 유실;
|
|
- reconnect replay 없음;
|
|
- Cluster node별 subscription semantics;
|
|
- expiry event가 TTL 0 시각에 정확히 오지 않음.
|
|
|
|
따라서 L1 flush/refresh hint로만 사용하고 hard TTL, generation, schema version을 recovery
|
|
boundary로 유지한다.
|
|
|
|
### 17.8 delete storm
|
|
|
|
mass invalidation 직후 모든 pod가 source를 동시에 load할 수 있다.
|
|
|
|
- generation bump;
|
|
- randomized prewarm;
|
|
- source concurrency budget;
|
|
- local/distributed single-flight;
|
|
- stale grace;
|
|
- queue/backpressure;
|
|
- progressive rollout
|
|
|
|
을 조합한다. invalidation producer가 모든 key를 즉시 delete하는 방식은 기본이 아니다.
|
|
|
|
## 18. Stampede와 source protection
|
|
|
|
### 18.1 방어 계층
|
|
|
|
권장 순서:
|
|
|
|
1. positive/negative bounded TTL;
|
|
2. TTL jitter;
|
|
3. local single-flight;
|
|
4. soft TTL + stale serve;
|
|
5. probabilistic early refresh;
|
|
6. distributed refresh lease;
|
|
7. source bulkhead;
|
|
8. load shedding.
|
|
|
|
분산 lock 하나로 stampede 전체를 해결하지 않는다.
|
|
|
|
### 18.2 TTL jitter
|
|
|
|
동일 batch로 생성된 entry가 같은 시점에 만료되지 않게 actual TTL을 bounded range에서 정한다.
|
|
|
|
```text
|
|
actual = configured * (1 + sample[-jitter, +jitter])
|
|
```
|
|
|
|
규칙:
|
|
|
|
- cryptographic randomness 불필요;
|
|
- hard minimum 보장;
|
|
- policy revision에 jitter strategy 기록;
|
|
- test에서는 seeded source로 deterministic 검증;
|
|
- compliance/authorization expiry를 늦춰서는 안 됨.
|
|
|
|
### 18.3 local single-flight
|
|
|
|
같은 process의 동시 miss는 한 loader future를 공유한다.
|
|
|
|
필수 bound:
|
|
|
|
- maximum in-flight keys;
|
|
- per-load deadline;
|
|
- waiter limit;
|
|
- completed entry 즉시 제거;
|
|
- cancellation semantics;
|
|
- loader exception fan-out;
|
|
- abandoned future reaper;
|
|
- key digest만 diagnostic에 사용.
|
|
|
|
single-flight map이 cache처럼 무한히 남지 않는다.
|
|
|
|
### 18.4 distributed refresh lease
|
|
|
|
다중 pod에서 한 owner만 refresh를 시도하도록 짧은 efficiency lease를 쓸 수 있다.
|
|
|
|
```text
|
|
claim refresh(key, owner, leaseTtl)
|
|
acquired -> source load -> captured generation put -> owner-safe release
|
|
contended -> stale serve or bounded wait
|
|
```
|
|
|
|
보장:
|
|
|
|
- 같은 Redis primary가 정상인 동안 duplicate load 감소;
|
|
- owner-safe release;
|
|
- lease TTL로 crashed loader 회복.
|
|
|
|
비보장:
|
|
|
|
- failover/partition에서 전역 단일 loader;
|
|
- source side effect correctness;
|
|
- loader completion 전 lease 유지.
|
|
|
|
loader는 read-only/idempotent여야 한다. lease를 잃어 duplicate load가 발생해도 business side
|
|
effect가 생기지 않아야 한다.
|
|
|
|
### 18.5 double check
|
|
|
|
distributed refresh lease를 획득한 뒤 cache를 다시 읽는다. 다른 owner가 먼저 채웠을 수 있다.
|
|
|
|
```text
|
|
miss -> claim -> lookup again -> still miss면 load
|
|
```
|
|
|
|
두 번째 lookup을 생략하면 불필요한 source load가 생긴다.
|
|
|
|
### 18.6 lease TTL
|
|
|
|
lease TTL은 source load deadline보다 길고 shutdown/GC pause risk를 고려하지만 무한하지 않다.
|
|
load duration distribution을 관측해 설정한다.
|
|
|
|
watchdog renewal은 baseline cache refresh에 필수로 두지 않는다. refresh가 TTL보다 길면:
|
|
|
|
- load를 cancel;
|
|
- duplicate load를 허용;
|
|
- source query를 paging/background job으로 바꿈
|
|
|
|
중 하나를 선택한다.
|
|
|
|
### 18.7 source fallback budget
|
|
|
|
Redis outage에서 모든 request를 DB로 보내면 cache failure가 DB outage로 확대된다.
|
|
|
|
region은 다음을 선언한다.
|
|
|
|
```text
|
|
maximum concurrent source loads
|
|
maximum queued waiters
|
|
load deadline
|
|
overload outcome
|
|
stale fallback
|
|
```
|
|
|
|
cache fail-open은 unlimited fail-open이 아니다.
|
|
|
|
### 18.8 hot key
|
|
|
|
한 key가 한 Redis shard와 한 source row에 집중될 수 있다.
|
|
|
|
- request coalescing;
|
|
- stale serve;
|
|
- refresh ahead;
|
|
- read replica/cache replica;
|
|
- payload split 금지 여부;
|
|
- local L1;
|
|
- hot-key metric/sample
|
|
|
|
을 검토한다. hash tag를 바꿔 동일 logical key를 여러 shard에 복제하면 invalidation과 consistency
|
|
cost가 늘어나므로 명시적 replicated-cache strategy일 때만 허용한다.
|
|
|
|
## 19. Edge rate-limit contract
|
|
|
|
### 19.1 소유권
|
|
|
|
transport edge rate limit은 `shared-contract`가 framework-neutral value contract를 소유한다.
|
|
|
|
```java
|
|
public interface EdgeRateLimitPort {
|
|
RateLimitOutcome evaluate(RateLimitRequest request);
|
|
}
|
|
```
|
|
|
|
`adapter:inbound:web`가 다음을 수행한다.
|
|
|
|
- trusted client IP 해석;
|
|
- authenticated principal/tenant/API key 해석;
|
|
- normalized route/operation ID 선택;
|
|
- policy ID 선택;
|
|
- bootstrap이 주입한 framework-neutral `EdgeSubjectPseudonymizer`로 canonical subject를
|
|
versioned HMAC digest로 변환;
|
|
- HTTP response/header mapping.
|
|
|
|
`EdgeSubjectPseudonymizer` contract는 `shared-contract`, secret-backed 구현과 rotation은
|
|
`adapter:outbound:identifier` 또는 별도 approved provider, wiring은 `app-bootstrap`이 소유한다.
|
|
inbound는 HMAC secret을 직접 resolve하지 않고 Redis provider는 raw identity나 HTTP route를
|
|
직접 파싱하지 않는다.
|
|
|
|
### 19.2 request
|
|
|
|
```java
|
|
public record RateLimitRequest(
|
|
String policyId,
|
|
String subjectDigest,
|
|
long cost,
|
|
String evaluationId,
|
|
Instant callerDeadline) {}
|
|
```
|
|
|
|
- `policyId`: bounded allowlisted ID;
|
|
- `subjectDigest`: inbound가 만든 versioned HMAC digest;
|
|
- `cost`: positive bounded integer;
|
|
- `evaluationId`: optional retry dedup token;
|
|
- `callerDeadline`: transport deadline budget.
|
|
|
|
algorithm, Redis key, window timestamp는 request에 넣지 않는다. provider의 policy registry가
|
|
소유한다.
|
|
|
|
### 19.3 outcome과 decision
|
|
|
|
```java
|
|
public sealed interface RateLimitOutcome {
|
|
record Evaluated(RateLimitDecision decision) implements RateLimitOutcome {}
|
|
record Degraded(RateLimitDecision decision, DegradationReason reason)
|
|
implements RateLimitOutcome {}
|
|
record Unavailable(
|
|
String policyId, Duration retryAfter, FailureCategory category)
|
|
implements RateLimitOutcome {}
|
|
record Indeterminate(
|
|
String policyId, String evaluationId, Duration retryAfter)
|
|
implements RateLimitOutcome {}
|
|
}
|
|
```
|
|
|
|
`allowed`가 의미 있는 경우에만 decision이 존재한다.
|
|
|
|
```java
|
|
public record RateLimitDecision(
|
|
boolean allowed,
|
|
long limit,
|
|
long remaining,
|
|
Duration retryAfter,
|
|
Instant resetAt,
|
|
String policyId,
|
|
String policyRevision,
|
|
DecisionSource source,
|
|
DecisionCertainty certainty) {}
|
|
```
|
|
|
|
`DecisionSource`:
|
|
|
|
```text
|
|
GLOBAL_REDIS
|
|
LOCAL_EMERGENCY
|
|
FAIL_OPEN_POLICY
|
|
SHADOW
|
|
```
|
|
|
|
`DecisionCertainty`는 `CERTAIN | APPROXIMATE_ALGORITHM`만 가진다. mutation certainty가 없는
|
|
경우 decision을 만들지 않고 `Indeterminate`를 반환한다.
|
|
|
|
기본 HTTP mapping:
|
|
|
|
| Outcome | Mapping |
|
|
| --- | --- |
|
|
| `Evaluated(allowed=true)` | request 진행, signaling header |
|
|
| `Evaluated(allowed=false)` | `429`, decision의 `Retry-After` |
|
|
| `Degraded(allowed=true/false)` | 해당 allow/429 + internal degraded telemetry |
|
|
| `Unavailable` | `503`, outcome의 `Retry-After` |
|
|
| `Indeterminate` | strict/default `503`; policy가 local fallback을 성공하면 `Degraded`로 변환 |
|
|
|
|
security product가 unavailable을 429로 숨겨야 하면 named HTTP mapping policy와 contract test를
|
|
별도 둔다. current 고정 1초 mapping은 제거한다.
|
|
|
|
### 19.4 policy
|
|
|
|
```java
|
|
public record RateLimitPolicy(
|
|
String id,
|
|
String revision,
|
|
RateLimitAlgorithm algorithm,
|
|
SubjectDimensions dimensions,
|
|
RateParameters parameters,
|
|
FailurePolicy failurePolicy,
|
|
DedupPolicy dedupPolicy,
|
|
CardinalityBudget cardinalityBudget,
|
|
boolean shadow) {}
|
|
```
|
|
|
|
global algorithm 하나가 아니라 policy별로 선택한다.
|
|
|
|
### 19.5 subject dimension
|
|
|
|
가능한 dimension:
|
|
|
|
```text
|
|
global
|
|
tenant
|
|
principal
|
|
api-key
|
|
client-ip
|
|
route/operation
|
|
resource class
|
|
```
|
|
|
|
raw value는 key, log, metric에 넣지 않는다. authenticated principal에도 route/policy dimension을
|
|
포함해 현재의 global quota collision을 제거한다.
|
|
|
|
### 19.6 business quota
|
|
|
|
“한 고객이 하루에 export 100개 생성 가능”처럼 domain/application rule인 quota는 별도
|
|
`BusinessQuotaPort`와 use-case policy다. HTTP abuse rate limit과 공유하면 transport 우회,
|
|
batch consumer, gRPC 호출에서 rule이 사라진다.
|
|
|
|
## 20. Rate-limit algorithm catalog
|
|
|
|
### 20.1 공통 원칙
|
|
|
|
모든 algorithm은:
|
|
|
|
- Function/Lua 한 번으로 read/decide/write;
|
|
- Redis `TIME` 기반 server time;
|
|
- bounded integer millisecond/fixed-point arithmetic;
|
|
- state TTL;
|
|
- policy revision key;
|
|
- maximum cost/state validation;
|
|
- Cluster same-slot;
|
|
- typed result;
|
|
- concurrency property test
|
|
|
|
를 갖는다.
|
|
|
|
client clock은 response `resetAt` 표시 보조로만 사용한다. enforcement calculation은 pod clock
|
|
skew의 영향을 줄이기 위해 Redis server time을 사용한다.
|
|
|
|
program은 `TIME`을 한 번만 읽고 stored `lastObservedMillis/windowId`와 비교한다.
|
|
|
|
- `serverNow < lastObserved`: `effectiveNow=max(serverNow,lastObserved)`로 clamp;
|
|
- token bucket/GCRA/sliding state를 뒤로 이동하지 않음;
|
|
- fixed window는 last accepted window ID보다 작은 window로 회귀하지 않음;
|
|
- forward jump의 refill은 capacity에서 saturation;
|
|
- configured unsafe clock-step threshold 초과 시 `CLOCK_UNSAFE` outcome;
|
|
- strict policy는 fail closed, availability policy는 explicit degraded fallback.
|
|
|
|
clock clamp가 Redis lease의 wall-clock safety를 strong하게 만들지는 않는다.
|
|
|
|
### 20.2 fixed window
|
|
|
|
state:
|
|
|
|
```text
|
|
windowId -> consumed
|
|
```
|
|
|
|
atomic steps:
|
|
|
|
1. `windowId=floor(effectiveNow/windowMillis)` 계산;
|
|
2. window interval은 `[windowId*windowMillis, (windowId+1)*windowMillis)`;
|
|
3. `cost <= limit`과 overflow 검증;
|
|
4. `consumed + cost <= limit`일 때만 counter를 증가;
|
|
5. first accepted request에서 TTL을
|
|
`windowEnd-effectiveNow+cleanupGraceMillis`로 설정;
|
|
6. denied request는 baseline에서 counter를 소비하지 않음;
|
|
7. `remaining=max(0, limit-newConsumed)`;
|
|
8. deny의 `retryAfter=windowEnd-effectiveNow`, `resetAt=windowEnd`.
|
|
|
|
특성:
|
|
|
|
- O(1) state;
|
|
- 이해하기 쉬움;
|
|
- boundary 직전/직후에 두 window quota를 연속 사용 가능;
|
|
- global smoothness가 필요 없는 단순 protection에 적합.
|
|
|
|
현재 local fixed-window를 Redis로 옮기는 최소 migration algorithm이지만 모든 policy의 default는
|
|
아니다.
|
|
|
|
### 20.3 sliding-window log
|
|
|
|
state:
|
|
|
|
```text
|
|
sorted set(member=evaluationId-or-unique-token, score=serverMillis)
|
|
```
|
|
|
|
atomic steps:
|
|
|
|
1. interval은 `(effectiveNow-windowMillis, effectiveNow]`;
|
|
2. `score <= effectiveNow-windowMillis` member를 bounded trim;
|
|
3. current count/cost 계산;
|
|
4. 허용 시 member 추가;
|
|
5. key TTL 설정;
|
|
6. oldest member에서 retry/reset 계산.
|
|
|
|
특성:
|
|
|
|
- event-level 정확한 sliding window;
|
|
- insert는 O(log N), `ZREMRANGEBYSCORE` trim은 O(log N + M);
|
|
- `M`은 한 invocation에서 제거하는 event 수이므로 per-call trim bound와 incremental cleanup 필요;
|
|
- event 수만큼 memory;
|
|
- attacker/high-volume policy에서 expensive;
|
|
- maximum members와 maximum policy rate를 startup에서 제한.
|
|
|
|
exact-log v1 profile은 `cost=1`만 허용한다. denied request는 member를 추가하지 않는다.
|
|
`retryAfter=max(1, oldestAcceptedScore+windowMillis-effectiveNow)`, TTL은
|
|
`windowMillis+cleanupGraceMillis`다. cost가 1보다 크면 member-per-cost로 확장하지 않고 별도
|
|
bounded weighted-log revision을 설계하거나 다른 algorithm을 선택한다.
|
|
|
|
### 20.4 sliding-window counter
|
|
|
|
state:
|
|
|
|
```text
|
|
previousWindowCount
|
|
currentWindowCount
|
|
```
|
|
|
|
estimate:
|
|
|
|
```text
|
|
windowId = floor(effectiveNow/windowMillis)
|
|
elapsed = effectiveNow - windowId*windowMillis
|
|
SCALE = 1_000_000
|
|
previousWeight = ceil((windowMillis-elapsed) * SCALE / windowMillis)
|
|
weightedScaled = current*SCALE + previous*previousWeight
|
|
```
|
|
|
|
두 key 또는 한 hash를 사용하며 같은 slot이다.
|
|
|
|
특성:
|
|
|
|
- O(1) state;
|
|
- fixed window보다 boundary burst 완화;
|
|
- exact log가 아닌 근사치;
|
|
- conservative `ceil` rounding;
|
|
- accepted current/previous count가 각 limit 이하일 때 exact log와의 absolute error upper bound는
|
|
`previousWindowCount <= limit`;
|
|
- 일반 API의 production option.
|
|
|
|
allow iff `weightedScaled + cost*SCALE <= limit*SCALE`; denied request는 current count를
|
|
증가시키지 않는다. state는 one versioned hash에 current/previous window ID/count와
|
|
`lastObservedMillis`를 저장하고 TTL은 `2*windowMillis+cleanupGraceMillis`다.
|
|
|
|
remaining을 exact quota처럼 표시하지 않고 `APPROXIMATE_ALGORITHM` certainty를 반환한다.
|
|
|
|
### 20.5 token bucket
|
|
|
|
state:
|
|
|
|
```text
|
|
tokensFixedPoint
|
|
lastRefillMillis
|
|
```
|
|
|
|
parameters:
|
|
|
|
```text
|
|
capacity
|
|
refillTokens
|
|
refillPeriod
|
|
requestCost
|
|
```
|
|
|
|
atomic steps:
|
|
|
|
1. `SCALE=1_000_000` micro-token으로 capacity/refill/cost 변환;
|
|
2. `elapsed=max(0,effectiveNow-lastRefillMillis)`;
|
|
3. `refill=floor(elapsed*refillScaled/refillPeriodMillis)`, multiply overflow 선검증;
|
|
4. `available=min(capacityScaled, storedTokens+refill)`;
|
|
5. `available>=costScaled`이면 차감, 아니면 state token을 차감하지 않음;
|
|
6. deny의
|
|
`retryAfter=ceil((costScaled-available)*refillPeriodMillis/refillScaled)`;
|
|
7. `resetAt`은 bucket full 시각,
|
|
`ceil((capacityScaled-newTokens)*refillPeriodMillis/refillScaled)`;
|
|
8. TTL은
|
|
`ceil(capacityScaled*refillPeriodMillis/refillScaled)+cleanupGraceMillis`.
|
|
|
|
특성:
|
|
|
|
- average rate와 burst capacity를 독립 제어;
|
|
- O(1) state;
|
|
- burst를 허용하는 API에 권장;
|
|
- floating point 대신 bounded fixed-point integer 사용;
|
|
- long idle 뒤 overflow를 막는 saturation arithmetic 필요.
|
|
|
|
정책 요구가 명확하지 않으면 “token bucket이 무조건 최고”로 고정하지 않는다.
|
|
|
|
### 20.6 leaky bucket
|
|
|
|
두 의미를 구분한다.
|
|
|
|
`policing`:
|
|
|
|
- 일정 rate를 넘는 요청을 즉시 reject;
|
|
- compact state로 구현 가능.
|
|
|
|
`shaping`:
|
|
|
|
- 허용 실행 시각을 계산해 queue에서 지연;
|
|
- synchronous HTTP request를 Redis 안이나 servlet thread에서 대기시키지 않음;
|
|
- background workflow/dispatcher가 bounded queue와 deadline을 소유할 때만 사용.
|
|
|
|
### 20.7 GCRA
|
|
|
|
state:
|
|
|
|
```text
|
|
theoreticalArrivalTime
|
|
```
|
|
|
|
장점:
|
|
|
|
- compact O(1) state;
|
|
- smooth quota;
|
|
- burst tolerance 표현.
|
|
|
|
위험:
|
|
|
|
- arithmetic/rounding 이해가 어렵고 operator 설명 비용이 큼;
|
|
- retry/reset 의미가 policy와 정확히 맞아야 함.
|
|
|
|
advanced opt-in으로 제공하며 token bucket과 동일 결과가 아님을 contract test로 고정한다.
|
|
|
|
### 20.8 concurrency limiter
|
|
|
|
동시에 실행 중인 request 수를 제한하는 것은 rate limit이 아니다.
|
|
|
|
별도 `ConcurrencyPermitPort`:
|
|
|
|
```text
|
|
acquire(subject, ttl)
|
|
renew(owner)
|
|
release(owner)
|
|
```
|
|
|
|
를 사용한다. permit leak, owner-safe release, lease expiry, queue bound를 다룬다. token bucket
|
|
cost로 concurrency를 흉내 내지 않는다.
|
|
|
|
### 20.9 algorithm comparison
|
|
|
|
| Algorithm | State | 정확성/특성 | 권장 |
|
|
| --- | --- | --- | --- |
|
|
| Fixed window | O(1) | boundary burst | simple protection |
|
|
| Sliding log | O(events) | exact sliding | low-volume high-value |
|
|
| Sliding counter | O(1) | bounded approximation | general API |
|
|
| Token bucket | O(1) | average + burst | burst-tolerant API |
|
|
| Leaky policing | O(1) | smooth rejection | no-burst policy |
|
|
| GCRA | O(1) | precise scheduling model | advanced |
|
|
|
|
### 20.10 normative vectors와 readiness
|
|
|
|
algorithm manifest에는 exact golden vector가 들어간다.
|
|
|
|
| Algorithm | Input/state | Expected |
|
|
| --- | --- | --- |
|
|
| fixed | window 1000ms, limit 2, at 999ms accepted=1, cost=1 | allow, remaining 0, reset 1000ms |
|
|
| fixed | same state at 999ms, cost=1 | deny, counter unchanged, retry 1ms |
|
|
| fixed | new request at 1000ms | new window, allow |
|
|
| sliding counter | previous=10, current=0, elapsed=500/1000ms | weighted=5 with configured scale/ceil |
|
|
| sliding counter | backward clock | window ID/state never regress, `CLOCK_UNSAFE` if threshold exceeded |
|
|
| token bucket | capacity 10, tokens 0, refill 10/1000ms, elapsed 250ms | 2.5 scaled tokens before cost |
|
|
| token bucket | available < cost | deny, token balance not deducted, exact ceil retry |
|
|
| sliding log | event score exactly `now-window` | trimmed; interval lower bound exclusive |
|
|
|
|
R2 baseline algorithms are fixed window, sliding counter, token bucket. Sliding log, GCRA, leaky
|
|
policing/shaping remain advanced until their own manifest, vector, state-growth and topology evidence
|
|
card passes. Test reference implementation uses the formulas above, not an independently guessed
|
|
algorithm.
|
|
|
|
## 21. Rate-limit policy composition과 failure
|
|
|
|
### 21.1 hierarchical policy
|
|
|
|
한 request에 global + tenant + principal + route limit이 동시에 적용될 수 있다.
|
|
|
|
선택지는:
|
|
|
|
1. 같은 slot의 bounded composite program으로 all-or-nothing evaluate;
|
|
2. 독립 policy를 순서대로 evaluate;
|
|
3. approximate/local upper-tier와 exact lower-tier 조합.
|
|
|
|
서로 다른 slot의 evaluation을 atomic하다고 표현하지 않는다.
|
|
|
|
순차 evaluate에서 앞 policy token을 소비한 뒤 뒤 policy가 deny할 수 있다. refund는 또 다른 race를
|
|
만든다. 이 conservative consumption을 명시하거나 same-slot composite를 사용한다.
|
|
|
|
### 21.2 hot global key
|
|
|
|
global policy 하나는 모든 traffic이 한 key/slot에 모인다. 다음을 검토한다.
|
|
|
|
- ingress/gateway 상위 limiter;
|
|
- shard별 approximate pre-limit;
|
|
- tenant/route partition;
|
|
- local emergency ceiling;
|
|
- dedicated rate-limit deployment;
|
|
- actual command latency/capacity evidence.
|
|
|
|
global exactness를 위해 한 hot key를 무한 확장할 수 있다고 가정하지 않는다.
|
|
|
|
### 21.3 evaluation dedup
|
|
|
|
response를 잃고 동일 request가 재시도되면 token이 두 번 차감될 수 있다.
|
|
|
|
strict cost boundary는 optional `evaluationId` dedup을 사용한다.
|
|
|
|
```text
|
|
evaluationId -> prior decision, short TTL
|
|
```
|
|
|
|
dedup record와 algorithm state는 같은 slot/program에서 처리한다. memory cost가 있으므로 policy별
|
|
enable, maximum IDs, TTL을 둔다.
|
|
|
|
dedup이 꺼져 있으면 at-least-once evaluation과 possible double charge를 descriptor에 명시한다.
|
|
|
|
### 21.4 failure modes
|
|
|
|
| Policy | Redis failure |
|
|
| --- | --- |
|
|
| abuse/security boundary | fail closed 또는 bounded local deny-first |
|
|
| monetary/cost protection | fail closed |
|
|
| general availability throttle | bounded local emergency limiter |
|
|
| non-critical smoothing | explicit fail open |
|
|
| shadow policy | allow + telemetry |
|
|
|
|
global `FAIL_OPEN=true`는 없다.
|
|
|
|
### 21.5 local emergency limiter
|
|
|
|
Redis unavailable일 때 선택 가능한 fallback:
|
|
|
|
- process-local;
|
|
- global quota보다 conservative;
|
|
- bounded maximum keys/weight;
|
|
- short TTL;
|
|
- no persistence;
|
|
- `LOCAL_EMERGENCY` decision;
|
|
- Redis recovery 후 자동 drain;
|
|
- pod 수에 따라 global exactness가 없음을 명시.
|
|
|
|
fallback map도 current local implementation처럼 unbounded면 안 된다.
|
|
|
|
primary/fallback ownership:
|
|
|
|
- `shared-contract`: `EdgeRateLimitPort`, outcome, provider-neutral fallback policy value;
|
|
- Redis leaf: `provider=redis` primary;
|
|
- inbound web: bounded `local-emergency` provider와 HTTP enforcement;
|
|
- app-bootstrap: primary와 optional degraded provider를 explicit selection으로 조립하는 composite;
|
|
- application/domain: transport quota fallback 없음.
|
|
|
|
`degraded-provider=local-emergency`일 때만 local map/sweeper/metric bean을 만든다. Redis
|
|
`Unavailable` 또는 reconcilable timeout만 composite fallback 후보이며 codec/program/config
|
|
failure에는 fallback하지 않는다. local result는 항상 `Degraded(source=LOCAL_EMERGENCY)`이고
|
|
global exactness를 광고하지 않는다.
|
|
|
|
`perPodLimit=floor(globalLimit * perPodShare)`로 capacity/refill/window limit을 보수적으로 줄인다.
|
|
결과가 0이면 해당 policy는 local allow를 하지 않고 fail closed한다.
|
|
`perPodShare * assumedMaximumPods <= 1`을 검증하지만
|
|
실제 pod가 가정을 초과하거나 traffic이 불균등하면 global quota가 아님을 descriptor/alert에
|
|
남긴다. maximum entries, entry TTL, in-flight, cleanup work도 설정 bound를 초과하지 않는다.
|
|
|
|
### 21.6 timeout certainty
|
|
|
|
rate program timeout 뒤 차감 여부가 indeterminate일 수 있다.
|
|
|
|
- evaluation dedup enabled: 같은 ID로 inspect/retry;
|
|
- strict policy without dedup: deny 또는 retry-after;
|
|
- availability policy: local emergency decision;
|
|
- 절대로 timeout을 ordinary allow로 조용히 바꾸지 않음.
|
|
|
|
### 21.7 shadow mode
|
|
|
|
정책 migration은 실제 deny 없이 decision을 기록하는 shadow mode를 지원한다.
|
|
|
|
- allowed response;
|
|
- would-have-denied metric;
|
|
- no subject metric tag;
|
|
- bounded sample log;
|
|
- state cost는 실제와 동일하므로 capacity 고려;
|
|
- shadow가 security control로 오인되지 않게 descriptor 표시.
|
|
|
|
## 22. Efficiency lease contract
|
|
|
|
### 22.1 기존 port의 위치
|
|
|
|
현재 `DistributedLockPort`는 “efficiency lock, DB constraint가 correctness authority”라는 문서가
|
|
있다. 이 의미는 유지한다. 기존 `tryAcquire(...)->DistributedLock.close()`는 compatibility
|
|
facade로 두고 새 v2 contract로 구현한다.
|
|
|
|
### 22.2 v2 request/outcome
|
|
|
|
```java
|
|
public interface DistributedLeasePort {
|
|
LeaseAttempt newAttempt(String operationId);
|
|
LeaseAcquireOutcome tryAcquire(LeaseRequest request);
|
|
LeaseInspectionOutcome inspect(LeaseInspectionRequest request);
|
|
}
|
|
```
|
|
|
|
```java
|
|
public record LeaseAttempt(String ownerToken, String operationId) {}
|
|
|
|
public record LeaseRequest(
|
|
String purpose,
|
|
String resourceDigest,
|
|
Duration waitTimeout,
|
|
Duration leaseTtl,
|
|
LeaseAttempt attempt) {}
|
|
|
|
public record LeaseInspectionRequest(
|
|
String purpose, String resourceDigest, LeaseAttempt attempt) {}
|
|
```
|
|
|
|
outcome:
|
|
|
|
```text
|
|
Acquired(LeaseHandle)
|
|
ReplayedSameOperation(LeaseHandle)
|
|
Contended(retryAfter)
|
|
OwnerOperationConflict
|
|
Unavailable(category)
|
|
Overloaded
|
|
Indeterminate(operationId)
|
|
|
|
Inspection:
|
|
Owned(LeaseHandle) | Absent | NotOwner | OwnerOperationConflict |
|
|
Unavailable | Indeterminate
|
|
```
|
|
|
|
`newAttempt`는 network/Redis side effect 없이 secure random opaque owner token을 만든다. caller는
|
|
최초 send 전에 반환된 attempt를 보관하고 retry/inspect에 같은 값을 사용한다. adapter 내부에서
|
|
send 직전에 token을 만들어 caller에게 숨기는 구현은 금지한다.
|
|
|
|
### 22.3 lease handle
|
|
|
|
```java
|
|
public interface LeaseHandle extends AutoCloseable {
|
|
String ownerToken();
|
|
String operationId();
|
|
Instant acquiredAt();
|
|
Duration remainingValidity();
|
|
boolean isUsableFor(Duration workBudget);
|
|
Instant observedServerExpiry(); // telemetry only
|
|
LeaseState state(); // ACTIVE | LOST | RELEASED | UNKNOWN
|
|
LeaseRenewOutcome renew();
|
|
LeaseReleaseOutcome release();
|
|
}
|
|
```
|
|
|
|
application이 Redis key를 보지 않는다. owner token은 secure random opaque value이며 log/metric에
|
|
남기지 않는다.
|
|
|
|
### 22.4 acquire
|
|
|
|
single-primary baseline은 owner와 operation ID를 한 bounded envelope에 저장한다.
|
|
|
|
```text
|
|
SET leaseKey ownerOperationEnvelope NX PX leaseTtl
|
|
```
|
|
|
|
same-attempt replay와 inspect가 필요한 R2 provider는 `lease-acquire-v1`/`lease-inspect-v1`
|
|
program으로 owner+operation을 비교한다. finite wait는 client에서 bounded backoff+jitter로
|
|
반복하며 한 Redis script가 wait하지 않는다.
|
|
|
|
### 22.5 release
|
|
|
|
금지:
|
|
|
|
```text
|
|
DEL leaseKey
|
|
```
|
|
|
|
old owner lease가 만료된 뒤 new owner가 acquire했을 수 있다.
|
|
|
|
필수:
|
|
|
|
```text
|
|
if GET leaseKey == ownerToken then DEL leaseKey
|
|
```
|
|
|
|
를 one atomic program으로 실행한다.
|
|
|
|
release outcome:
|
|
|
|
```text
|
|
RELEASED
|
|
ALREADY_ABSENT
|
|
NOT_OWNER
|
|
INDETERMINATE
|
|
UNAVAILABLE
|
|
```
|
|
|
|
기존 `void close()` compatibility facade는 release result를 caller에게 전달할 수 없다.
|
|
따라서 `NOT_OWNER/INDETERMINATE`는 telemetry와 lease-lost callback에만 남기고, 결과에 따라
|
|
application policy를 실행해야 하는 consumer는 반드시 v2 `release()` outcome으로 migration한다.
|
|
legacy facade가 silent success를 보장한다고 문서화하지 않는다.
|
|
|
|
### 22.6 renew
|
|
|
|
renew도 owner compare 후 TTL을 바꾼다.
|
|
|
|
```text
|
|
if GET leaseKey == ownerToken then PEXPIRE leaseKey newTtl
|
|
```
|
|
|
|
renew timeout은 lease가 연장되었는지 알 수 없는 `INDETERMINATE`다. correctness-sensitive work는
|
|
즉시 lease를 `UNKNOWN/LOST`로 보고 protected operation을 중단해야 한다.
|
|
|
|
### 22.7 watchdog
|
|
|
|
watchdog를 사용할 경우:
|
|
|
|
- fixed cadence가 lease TTL보다 충분히 짧음;
|
|
- scheduling delay/GC pause 관측;
|
|
- renewal queue bounded;
|
|
- application deadline 이후 renew 금지;
|
|
- shutdown 시 new renew 중단;
|
|
- consecutive failure threshold가 아니라 validity deadline으로 lost 판단;
|
|
- handle state thread-safe.
|
|
|
|
watchdog가 process pause나 failover를 제거하지 않는다.
|
|
|
|
### 22.8 validity
|
|
|
|
acquire response latency를 뺀 effective validity를 계산한다.
|
|
|
|
```text
|
|
remainingValidity =
|
|
leaseTtl - localMonotonicElapsedSinceAcquireStart - driftBudget
|
|
```
|
|
|
|
remaining이 minimum protected-work budget보다 작으면 acquired result를 사용하지 않고 release한다.
|
|
`Instant` wall-clock은 authoritative validity 판단에 사용하지 않는다. `observedServerExpiry`는
|
|
operator telemetry일 뿐이다. Redis clock step, failover, renewal timeout을 감지하면 wall-clock
|
|
추정과 무관하게 handle을 `UNKNOWN/LOST`로 전환한다.
|
|
|
|
### 22.9 unknown acquire
|
|
|
|
acquire command는 적용되었는데 response를 잃을 수 있다. caller가 보관한 같은
|
|
`LeaseAttempt(ownerToken, operationId)`로 inspect하거나 acquire를 반복한다. live envelope가
|
|
일치하면 같은 handle/remaining TTL을 `Owned`/`ReplayedSameOperation`으로 회수한다. owner는 같고
|
|
operation이 다르면 conflict이며, absent 또는 inspect도 timeout이면 이전 acquire를 성공/실패로
|
|
단정하지 않는다.
|
|
|
|
random new token으로 즉시 재시도하면 self-contention이나 two-attempt confusion이 생긴다.
|
|
|
|
### 22.10 사용 가능 범위
|
|
|
|
적합:
|
|
|
|
- duplicate cache refresh 감소;
|
|
- duplicate scheduled cleanup 감소;
|
|
- cost가 낮고 idempotent한 background work;
|
|
- DB constraint가 최종 authority인 mutation의 contention 완화.
|
|
|
|
부적합:
|
|
|
|
- 결제 중복 방지의 유일한 장치;
|
|
- inventory invariant;
|
|
- unique ID authority;
|
|
- external device exclusive command;
|
|
- stale writer를 거절할 수 없는 storage write.
|
|
|
|
## 23. Fencing과 coordination capability
|
|
|
|
### 23.1 별도 contract
|
|
|
|
fencing은 efficiency lease의 boolean option이 아니다.
|
|
|
|
```java
|
|
public interface FencedLeasePort {
|
|
FencedLeaseAttempt newAttempt(String operationId);
|
|
FencedLeaseAcquireOutcome tryAcquire(FencedLeaseRequest request);
|
|
FencedLeaseInspectionOutcome inspect(FencedLeaseInspectionRequest request);
|
|
FencedLeaseRenewOutcome renew(FencedLeaseHandle handle, Duration leaseTtl);
|
|
FencedLeaseReleaseOutcome release(FencedLeaseHandle handle);
|
|
}
|
|
```
|
|
|
|
`FencedLeaseRequest`와 inspection request는 caller가 최초 send 전 받은 같은
|
|
`FencedLeaseAttempt(ownerToken, operationId)`와 durable
|
|
`FencingResourceRegistration(resourceEpoch, registrationDigest)`를 포함한다. response-loss retry
|
|
중 adapter가 새 owner token/epoch를 만들지 않는다.
|
|
|
|
handle:
|
|
|
|
```text
|
|
ownerToken
|
|
operationId
|
|
fencingToken(resourceEpoch, counter)
|
|
remainingValidity/isUsableFor
|
|
renew/release/lost state
|
|
```
|
|
|
|
acquire outcome:
|
|
|
|
```text
|
|
ACQUIRED(handle)
|
|
REPLAYED_SAME_OPERATION(handle)
|
|
CONTENDED(retryAfter)
|
|
OWNER_OPERATION_CONFLICT
|
|
FENCE_COUNTER_MISSING
|
|
FENCE_REGRESSION
|
|
FENCE_EXHAUSTED
|
|
EPOCH_MISMATCH
|
|
REGISTRATION_CONFLICT
|
|
UNAVAILABLE_BEFORE_SEND
|
|
INDETERMINATE(operationId)
|
|
|
|
Inspection:
|
|
OWNED(handle) | ABSENT | NOT_OWNER | OWNER_OPERATION_CONFLICT |
|
|
FENCE_COUNTER_MISSING | EPOCH_MISMATCH | REGISTRATION_CONFLICT |
|
|
FENCE_REGRESSION | UNAVAILABLE | INDETERMINATE
|
|
Renew:
|
|
RENEWED | ABSENT | NOT_OWNER | TOKEN_MISMATCH | FENCE_COUNTER_MISSING |
|
|
EPOCH_MISMATCH | REGISTRATION_CONFLICT | FENCE_REGRESSION |
|
|
INDETERMINATE | UNAVAILABLE
|
|
Release:
|
|
RELEASED | ABSENT | NOT_OWNER | TOKEN_MISMATCH | FENCE_COUNTER_MISSING |
|
|
EPOCH_MISMATCH | REGISTRATION_CONFLICT | FENCE_REGRESSION |
|
|
INDETERMINATE | UNAVAILABLE
|
|
```
|
|
|
|
`Instant validUntil`은 authority가 아니며 §22와 같은 local monotonic budget을 사용한다.
|
|
inspection은 resource/owner/operation ID가 모두 같은 live owner record일 때만 기존 fencing
|
|
token과 remaining TTL을 돌려준다.
|
|
|
|
fenced owner envelope는 일반 lease envelope와 schema/epoch/counter가 다르므로 generic
|
|
`lease-renew/release` program을 재사용하지 않는다. §13.10의 fenced-specific renew/release가
|
|
handle의 epoch/counter와 current registration까지 비교한다. renew timeout은 handle을
|
|
`UNKNOWN/LOST`로 만들고 자동 재시도하지 않으며, release timeout은 desired-absent repeat/inspect
|
|
전까지 `INDETERMINATE`다.
|
|
|
|
### 23.2 protected resource requirement
|
|
|
|
fencing token은 lock provider가 발급하는 것만으로 충분하지 않다. protected resource가 마지막
|
|
accepted token을 저장하고:
|
|
|
|
```text
|
|
incomingToken.epoch < lastAcceptedToken.epoch -> reject
|
|
incomingToken.epoch == lastAcceptedToken.epoch
|
|
&& incomingToken.counter <= lastAcceptedToken.counter -> reject
|
|
```
|
|
|
|
해야 한다. epoch가 더 큰 token은 protected resource의 durable registration/activation과
|
|
일치할 때만 받아들이며 Redis caller가 임의 epoch를 높일 수 없다.
|
|
|
|
resource가 token을 검증할 수 없으면 `FENCED` guarantee를 광고하지 않는다.
|
|
|
|
### 23.3 Redis counter failover
|
|
|
|
fencing counter도 Redis 비동기 replication에서 acknowledged increment가 유실될 수 있다. promoted
|
|
replica가 더 낮은 token을 발급할 수 있다.
|
|
|
|
이미 높은 token을 본 protected resource가 낮은 token을 거절하면 stale safety는 유지될 수 있지만,
|
|
counter가 high watermark를 넘어갈 때까지 새 work도 거절되어 availability가 떨어진다.
|
|
|
|
따라서:
|
|
|
|
- counter persistence/replication profile 명시;
|
|
- resource-side high watermark;
|
|
- token regression alert;
|
|
- recovery runbook;
|
|
- “Redis counter이므로 monotonic forever” 문구 금지.
|
|
|
|
### 23.3.1 fenced acquire program과 counter lifecycle
|
|
|
|
`fenced-lease-acquire`는 같은 slot의 두 key를 한 program에서 처리한다.
|
|
|
|
```text
|
|
KEYS[1] = lease owner key
|
|
KEYS[2] = fence counter key
|
|
ARGV = expectedResourceEpoch, registrationDigest, ownerToken,
|
|
leaseTtlMillis, operationId, hardFailThreshold
|
|
```
|
|
|
|
first write 전에 key type, current owner, TTL, counter integer/range, provisioned epoch/digest를
|
|
모두 검증한다.
|
|
|
|
- live owner token과 operation ID가 모두 같으면 counter를 증가시키지 않고 기존 fencing token을
|
|
`REPLAYED_SAME_OPERATION`으로 반환한다.
|
|
- owner token은 같지만 operation ID가 다르면 `OWNER_OPERATION_CONFLICT`다.
|
|
- 다른 live owner면 `CONTENDED`다.
|
|
- owner가 없을 때만:
|
|
|
|
1. signed 64-bit counter를 1 증가;
|
|
2. new `(resourceEpoch, counter)` fencing token을 얻음;
|
|
3. owner token + epoch/counter + operation ID를 lease TTL과 함께 기록
|
|
|
|
한다.
|
|
|
|
acquire 응답 유실 뒤 caller는 같은 owner/operation ID로 `inspect`하거나 동일 acquire를 반복한다.
|
|
live record가 남아 있으면 같은 fencing token을 회수하고, 이미 만료되었으면 새 operation으로
|
|
정책상 재시도하되 이전 effect를 자동 성공/실패로 단정하지 않는다. inspect/replay에도 실패하면
|
|
`INDETERMINATE`를 유지한다.
|
|
|
|
counter 규칙:
|
|
|
|
- coordination `noeviction` role;
|
|
- TTL 없음;
|
|
- resource retirement 없이 cleanup 금지;
|
|
- missing/regressed counter를 0으로 초기화하지 않음;
|
|
- request epoch/digest와 provisioned envelope가 다르면 `EPOCH_MISMATCH`/
|
|
`REGISTRATION_CONFLICT`;
|
|
- `FENCE_REGRESSION/UNAVAILABLE`로 fail closed;
|
|
- signed 64-bit overflow 이전 configured hard-fail threshold;
|
|
- protected-resource high watermark와 operator recovery 필수.
|
|
|
|
#### Counter registration과 restore
|
|
|
|
missing counter는 “신규 resource”와 “Redis loss/restore”를 구분할 수 없으므로 request path에서
|
|
`SET NX 0`으로 만들지 않는다. protected resource의 durable store가 다음 registration을
|
|
authority로 소유한다.
|
|
|
|
```text
|
|
resourceId
|
|
resourceEpoch
|
|
registrationDigest
|
|
fencingStatus = PENDING | ACTIVE | RETIRED
|
|
lastAcceptedHighWatermark
|
|
```
|
|
|
|
one-time provisioning protocol:
|
|
|
|
1. protected resource 생성 transaction에서 random/durable epoch와 registration digest를 만들고
|
|
`PENDING`, high watermark 0을 commit한다.
|
|
2. after-commit reconciler가 `fenced-counter-provision-v1`을 호출한다. program은
|
|
`(epoch, registrationDigest, durableHighWatermark)` envelope를 no-TTL/noeviction key에
|
|
`NX`로 만들며 같은 registration은 `ALREADY_SAME`, 다른 값은 conflict다.
|
|
3. Redis read-back receipt의 epoch/digest/counter를 검증한 뒤 durable row를 `ACTIVE`로 바꾼다.
|
|
4. `ACTIVE` 전에는 fenced acquire를 fail closed한다.
|
|
|
|
DB commit 뒤 Redis provisioning 전 crash는 PENDING reconciler가 복구한다. Redis counter가
|
|
missing인데 durable row가 `ACTIVE`이면 신규 resource로 재해석하지 않는다. reconciler/operator는
|
|
protected resource의 durable high watermark와 epoch를 읽고, backup/incident evidence가
|
|
충분할 때만 counter를 그 high watermark 이상으로 reprovision한다. 다음 acquire increment가
|
|
반드시 마지막 accepted token보다 커야 한다. authoritative high watermark를 얻을 수 없으면
|
|
availability를 닫은 채 복구하지 않는다.
|
|
|
|
rollback/cleanup:
|
|
|
|
- `PENDING`이며 protected write가 전혀 없을 때만 registration digest compare-delete 후 row rollback;
|
|
- `ACTIVE` counter는 application rollback이나 generic cache cleanup으로 삭제 금지;
|
|
- retirement는 durable `RETIRED` tombstone과 epoch를 남기고 모든 holder/work drain 및 retention
|
|
뒤 operator workflow로 정리;
|
|
- resource ID 재사용 시 이전 epoch를 재사용하지 않음.
|
|
|
|
resource epoch는 모든 fenced resource에서 항상 존재한다. cleanup은 epoch를 나중에 “추가”하는
|
|
절차가 아니라 위 registration/retirement lifecycle과 `(epoch,counter)` 비교를 계속 보존한다.
|
|
|
|
### 23.4 Redlock
|
|
|
|
multi-master Redlock을 기본 correctness provider로 선택하지 않는다.
|
|
|
|
이유:
|
|
|
|
- finite lease와 wall-clock/drift 가정;
|
|
- network partition/GC pause;
|
|
- quorum acquire response uncertainty;
|
|
- 각 master state cleanup;
|
|
- protected resource fencing 필요성은 여전히 남음.
|
|
|
|
특정 product가 Redlock을 선택하면 별도 ADR, failure model, clock assumptions, quorum topology,
|
|
fenced consumer test가 필요하다. template 기본 descriptor는 `STRICT_COORDINATION_REQUIRED`를
|
|
Redis로 충족하지 않는다.
|
|
|
|
### 23.5 leader election
|
|
|
|
leader election은 lease 위에 semantic contract로 제공한다.
|
|
|
|
- epoch/fencing token;
|
|
- lease-lost callback;
|
|
- leader-only task cancellation;
|
|
- takeover delay;
|
|
- no singleton business correctness claim;
|
|
- scheduler work idempotency.
|
|
|
|
bean name `outboxLeaderElection` 존재만으로 안전을 판단하지 않는다.
|
|
|
|
### 23.6 semaphore
|
|
|
|
bounded distributed semaphore는 owner token별 permit record와 TTL이 필요하다.
|
|
|
|
- maximum permits;
|
|
- owner-safe release;
|
|
- crashed owner expiry;
|
|
- renewal;
|
|
- list cleanup bound;
|
|
- fairness non-guarantee;
|
|
- Cluster same-slot;
|
|
- exact current count reconciliation.
|
|
|
|
large owner set를 한 Lua에서 전부 scan하지 않는다.
|
|
|
|
### 23.7 work claim
|
|
|
|
queue/job claim은 lock과 다른 contract다.
|
|
|
|
```text
|
|
claim item -> owner/attempt/lease
|
|
ack success
|
|
nack retry
|
|
reclaim expired
|
|
```
|
|
|
|
ordering, retry count, terminal state가 필요하면 Redis Streams나 durable DB queue의 semantic
|
|
contract를 사용한다. 단순 lease key로 queue를 만들지 않는다.
|
|
|
|
## 24. Redis idempotency design
|
|
|
|
### 24.1 guarantee 이름
|
|
|
|
Redis provider의 기본 guarantee는:
|
|
|
|
```text
|
|
REQUEST_REPLAY
|
|
```
|
|
|
|
이다. 다음을 뜻한다.
|
|
|
|
- live record가 유지되는 동안 같은 request fingerprint를 식별;
|
|
- completed response를 replay;
|
|
- concurrent duplicate에 one current owner를 선택;
|
|
- owner-safe transition.
|
|
|
|
다음을 뜻하지 않는다.
|
|
|
|
- JDBC business effect exactly-once;
|
|
- external API side effect exactly-once;
|
|
- failover에서도 record zero-loss;
|
|
- arbitrary long-term dedup.
|
|
|
|
### 24.2 v2 claim
|
|
|
|
기존 `find -> tryBegin`을 한 atomic operation으로 바꾼다.
|
|
|
|
```java
|
|
IdempotencyClaimOutcome claim(IdempotencyClaimRequest request);
|
|
```
|
|
|
|
outcome:
|
|
|
|
```text
|
|
ACQUIRED(ownerToken, attempt, processingLeaseUntil)
|
|
REPLAYED_ACQUIRE(ownerToken, attempt, processingLeaseUntil)
|
|
TAKEN_OVER_CLAIMED(ownerToken, attempt, processingLeaseUntil)
|
|
COMPLETED_REPLAY(storedResponse, replayUntil)
|
|
IN_PROGRESS(retryAfter, currentAttempt)
|
|
RECOVERY_REQUIRED(currentAttempt)
|
|
FINGERPRINT_MISMATCH
|
|
OWNER_OPERATION_CONFLICT
|
|
INDETERMINATE(operationId)
|
|
UNAVAILABLE
|
|
```
|
|
|
|
claim caller는 최초 Redis send 전에 local `newClaimAttempt(operationId)`로 secure random
|
|
owner token을 받고 request와 함께 보관한다. same scope/fingerprint/owner/operation의 duplicate
|
|
claim은 counter/attempt를 바꾸지 않고 `REPLAYED_ACQUIRE`로 동일 owner handle을 반환한다.
|
|
|
|
### 24.2.1 complete v2 port
|
|
|
|
claim만 바꾸고 기존 scope-only mutation을 남기지 않는다.
|
|
|
|
```java
|
|
public interface IdempotencyStorePortV2 {
|
|
IdempotencyClaimAttempt newClaimAttempt(String operationId);
|
|
IdempotencyClaimOutcome claim(IdempotencyClaimRequest request);
|
|
|
|
IdempotencyStartOutcome markExecutionStarted(
|
|
IdempotencyOwner owner, String operationId);
|
|
|
|
IdempotencyRenewOutcome renew(
|
|
IdempotencyOwner owner, Duration processingLeaseTtl, String operationId);
|
|
|
|
IdempotencyCompleteOutcome complete(
|
|
IdempotencyOwner owner,
|
|
StoredResponse response,
|
|
Duration replayTtl,
|
|
String operationId);
|
|
|
|
IdempotencyFailOutcome markFailed(
|
|
IdempotencyOwner owner,
|
|
IdempotencyFailureDisposition disposition,
|
|
Duration retention,
|
|
String operationId);
|
|
|
|
IdempotencyReleaseOutcome releaseBeforeExecution(
|
|
IdempotencyOwner owner, String operationId);
|
|
|
|
IdempotencyInspection inspect(IdempotencyInspectionRequest request);
|
|
}
|
|
|
|
public record IdempotencyClaimAttempt(String ownerToken, String operationId) {}
|
|
|
|
public record IdempotencyOwner(
|
|
IdempotencyScope scope, String ownerToken, long attempt) {}
|
|
|
|
public record IdempotencyInspectionRequest(
|
|
IdempotencyScope scope,
|
|
String requestFingerprint,
|
|
IdempotencyClaimAttempt attempt) {}
|
|
|
|
public sealed interface VerifiedIdempotencyReconciliationEvidence
|
|
permits VerifiedCommittedEvidence, VerifiedNoEffectEvidence {
|
|
String receiptDigest();
|
|
String evidenceType();
|
|
String evidenceRevision();
|
|
}
|
|
|
|
public sealed interface VerifiedCommittedEvidence
|
|
extends VerifiedIdempotencyReconciliationEvidence
|
|
permits SourceCommittedEvidence {}
|
|
|
|
public sealed interface VerifiedNoEffectEvidence
|
|
extends VerifiedIdempotencyReconciliationEvidence
|
|
permits SourceNoEffectEvidence {}
|
|
|
|
public interface IdempotencyEffectEvidenceVerifier {
|
|
CommittedEvidenceVerificationOutcome verifyCommitted(
|
|
IdempotencyReconciliationCandidate candidate);
|
|
NoEffectEvidenceVerificationOutcome verifyNoEffect(
|
|
IdempotencyReconciliationCandidate candidate);
|
|
}
|
|
|
|
public interface IdempotencyReconciliationPort {
|
|
IdempotencyReconcileCommittedOutcome reconcileCommitted(
|
|
IdempotencyCommittedReconciliation request);
|
|
|
|
IdempotencyReopenOutcome reconcileNoEffectAndReopen(
|
|
IdempotencyNoEffectReconciliation request);
|
|
}
|
|
|
|
public record IdempotencyReconciliationAudit(
|
|
String actorDigest,
|
|
String reasonCode,
|
|
String evidenceReferenceDigest,
|
|
Instant requestedAt) {}
|
|
|
|
public record IdempotencyCommittedReconciliation(
|
|
IdempotencyScope scope,
|
|
long expectedAttempt,
|
|
long expectedStateRevision,
|
|
VerifiedCommittedEvidence evidence,
|
|
IdempotencyReconciliationAudit audit,
|
|
String auditOperationId,
|
|
StoredResponse response,
|
|
Duration replayTtl) {}
|
|
|
|
public record IdempotencyNoEffectReconciliation(
|
|
IdempotencyScope scope,
|
|
long expectedAttempt,
|
|
long expectedStateRevision,
|
|
VerifiedNoEffectEvidence evidence,
|
|
IdempotencyReconciliationAudit audit,
|
|
String auditOperationId,
|
|
IdempotencyClaimAttempt newAttempt,
|
|
Duration processingLeaseTtl) {}
|
|
```
|
|
|
|
permitted evidence implementation의 constructor/factory는 application reconciliation package
|
|
내부이며 verifier 성공 결과만 생성한다. arbitrary controller DTO는 이 sealed value를 구현하거나
|
|
deserialize할 수 없다. committed request는 `VerifiedCommittedEvidence`, reopen request는
|
|
`VerifiedNoEffectEvidence`만 받아 evidence 방향을 type-level로 뒤집을 수 없게 한다.
|
|
|
|
typed outcome:
|
|
|
|
```text
|
|
Start:
|
|
STARTED | ALREADY_STARTED_SAME_OPERATION | ABSENT | NOT_OWNER |
|
|
NOT_CLAIMED | OPERATION_CONFLICT | INDETERMINATE | UNAVAILABLE
|
|
Renew:
|
|
RENEWED | ALREADY_RENEWED_SAME_OPERATION | ABSENT | NOT_OWNER |
|
|
NOT_IN_PROGRESS | OPERATION_CONFLICT | INDETERMINATE | UNAVAILABLE
|
|
Complete:
|
|
COMPLETED | ALREADY_COMPLETED_SAME_RESULT | RESPONSE_CONFLICT |
|
|
ABSENT | NOT_OWNER | NOT_IN_PROGRESS | OPERATION_CONFLICT |
|
|
INDETERMINATE | UNAVAILABLE
|
|
Fail:
|
|
MARKED_RETRYABLE | MARKED_ABANDONED | ALREADY_MARKED_SAME_OPERATION |
|
|
ABSENT | NOT_OWNER | NOT_IN_PROGRESS | OPERATION_CONFLICT |
|
|
INDETERMINATE | UNAVAILABLE
|
|
Release:
|
|
RELEASED_BEFORE_EXECUTION | ALREADY_RELEASED_SAME_OPERATION |
|
|
ABSENT | NOT_OWNER | EXECUTION_ALREADY_STARTED | OPERATION_CONFLICT |
|
|
INDETERMINATE | UNAVAILABLE
|
|
Inspect:
|
|
ABSENT | CLAIMED_SAME_OPERATION(owner,lease) |
|
|
EXECUTING_SAME_OPERATION(owner,lease) | COMPLETED_REPLAY(response) |
|
|
IN_PROGRESS_OTHER | FAILED_RETRYABLE | ABANDONED |
|
|
FINGERPRINT_MISMATCH | OPERATION_CONFLICT | UNAVAILABLE
|
|
Reconcile committed:
|
|
RECONCILED_COMPLETED | ALREADY_RECONCILED_SAME_OPERATION |
|
|
EVIDENCE_CONFLICT | STATE_CONFLICT | ABSENT | INDETERMINATE | UNAVAILABLE
|
|
Reopen no effect:
|
|
REOPENED_CLAIMED(owner,attempt,lease) | ALREADY_REOPENED_SAME_OPERATION |
|
|
EVIDENCE_CONFLICT | STATE_CONFLICT | ABSENT | INDETERMINATE | UNAVAILABLE
|
|
```
|
|
|
|
같은 transition kind와 `lastTransitionOperationId`의 duplicate는 prior result를 replay한다.
|
|
다른 operation ID가 이미 끝난 동일 transition을 바꾸려 하면 conflict다. claim replay는 별도
|
|
`claimOperationId`를 비교하므로 start/renew가 claim recovery 정보를 덮어쓰지 않는다. claim 응답
|
|
유실 뒤 caller는 보관한 request의 fingerprint/attempt로 `inspect`하거나 동일 claim을 반복해
|
|
owner/attempt를 회수한다.
|
|
inspect도 unavailable이면 claim은 `INDETERMINATE`이며 새 owner로 즉시 claim하지 않는다.
|
|
|
|
### 24.3 request
|
|
|
|
```text
|
|
scopeDigest
|
|
requestFingerprint
|
|
claimAttempt(ownerToken, operationId)
|
|
processingLeaseTtl
|
|
replayTtl
|
|
responseCodecId
|
|
policyRevision
|
|
```
|
|
|
|
scope raw principal/idempotency key는 adapter에 전달하기 전에 canonical digest로 바꿀 수 있다.
|
|
application value에는 provider key syntax가 없다.
|
|
|
|
### 24.4 record
|
|
|
|
```text
|
|
recordVersion
|
|
state = CLAIMED | EXECUTING | COMPLETED | FAILED_RETRYABLE | ABANDONED
|
|
stateRevision
|
|
requestFingerprint
|
|
ownerToken
|
|
attempt
|
|
claimOperationId
|
|
lastTransitionOperationId?
|
|
lastTransitionKind?
|
|
lastTransitionResultDigest?
|
|
reconciliationEvidenceDigest?
|
|
processingLeaseUntil
|
|
responseCodecId?
|
|
responseVersion?
|
|
responseDigest?
|
|
responsePayload?
|
|
replayUntil?
|
|
createdAt
|
|
updatedAt
|
|
policyRevision
|
|
```
|
|
|
|
processing lease와 completed replay TTL은 분리한다. 30초 execution lease 때문에 completed
|
|
response가 30초 후 사라지거나, 24시간 replay TTL 때문에 crashed owner가 24시간 request를
|
|
막으면 안 된다.
|
|
|
|
### 24.5 state machine
|
|
|
|
```mermaid
|
|
stateDiagram-v2
|
|
[*] --> CLAIMED: first claim
|
|
CLAIMED --> EXECUTING: markExecutionStarted
|
|
CLAIMED --> [*]: release before execution
|
|
CLAIMED --> CLAIMED: same owner renew
|
|
EXECUTING --> EXECUTING: same owner renew
|
|
CLAIMED --> CLAIMED: expired pre-execution owner takeover / attempt+1
|
|
EXECUTING --> ABANDONED: expired execution / effect unknown
|
|
EXECUTING --> COMPLETED: owner-safe complete
|
|
EXECUTING --> FAILED_RETRYABLE: no-effect confirmed
|
|
EXECUTING --> ABANDONED: effect unknown/reconciliation
|
|
FAILED_RETRYABLE --> CLAIMED: retry claim
|
|
ABANDONED --> CLAIMED: explicit reconciliation proves safe retry
|
|
ABANDONED --> COMPLETED: committed receipt reconciliation
|
|
COMPLETED --> [*]: replay TTL expires
|
|
FAILED_RETRYABLE --> [*]: retry record expires
|
|
ABANDONED --> [*]: audit TTL expires
|
|
```
|
|
|
|
### 24.5.1 executor와 action lifecycle
|
|
|
|
기존 `Supplier<R>`와 “모든 RuntimeException에서 discard” contract를 제거한다.
|
|
|
|
```java
|
|
public interface IdempotentAction<R> {
|
|
IdempotentActionOutcome<R> execute(IdempotencyOwner owner);
|
|
}
|
|
|
|
public sealed interface IdempotentActionOutcome<R> {
|
|
record Committed<R>(R result, EffectReceipt receipt)
|
|
implements IdempotentActionOutcome<R> {}
|
|
record NoEffectConfirmed<R>(SourceFailure failure)
|
|
implements IdempotentActionOutcome<R> {}
|
|
record EffectUnknown<R>(SourceFailure failure)
|
|
implements IdempotentActionOutcome<R> {}
|
|
record CancelledBeforeStart<R>() implements IdempotentActionOutcome<R> {}
|
|
}
|
|
```
|
|
|
|
`IdempotencyExecutorV2` transition:
|
|
|
|
1. claim/inspect로 owner handle을 확정한다.
|
|
2. action을 호출하기 직전에 `markExecutionStarted`를 실행한다.
|
|
3. start 결과가 `STARTED`/`ALREADY_STARTED_SAME_OPERATION`일 때만 action을 호출한다.
|
|
4. start가 indeterminate이면 inspect로 `EXECUTING_SAME_OPERATION`을 확인하기 전에는 action을
|
|
호출하지 않는다.
|
|
5. caller cancellation이 successful start보다 먼저면 `releaseBeforeExecution`; start 뒤면 action
|
|
outcome/transaction evidence로만 fail/abandon을 결정한다.
|
|
|
|
| Action outcome | Store transition |
|
|
| --- | --- |
|
|
| `Committed` | owner-safe `complete`; completion indeterminate면 result를 exactly-once라고 응답하지 않고 reconcile |
|
|
| `NoEffectConfirmed` | `FAILED_RETRYABLE` 또는 owner-safe release policy |
|
|
| `CancelledBeforeStart` | `releaseBeforeExecution` |
|
|
| `EffectUnknown` | `ABANDONED/INDETERMINATE`, automatic retry 금지 |
|
|
| unclassified thrown exception | default `EffectUnknown`, 절대 delete/release하지 않음 |
|
|
|
|
`EffectReceipt`는 domain operation ID, committed source revision, downstream idempotency receipt처럼
|
|
실제 effect를 reconcile할 bounded reference다. Redis key/SDK type은 없다.
|
|
|
|
`TransactionalIdempotentAction` helper는 `TransactionPort`와 연결한다.
|
|
|
|
```text
|
|
validation before transaction fails -> NOT_STARTED/NO_EFFECT_CONFIRMED
|
|
transaction rolls back and rollback is confirmed -> NO_EFFECT_CONFIRMED
|
|
transaction commit returns successfully -> COMMITTED
|
|
commit response/connection state unknown -> EFFECT_UNKNOWN
|
|
external side effect inside transaction callback -> 별도 provider receipt 없으면 EFFECT_UNKNOWN
|
|
```
|
|
|
|
Redis `complete`는 JDBC transaction 안에 넣어 atomic하다고 가장하지 않고 DB commit 뒤 실행한다.
|
|
same-store JPA provider가 claim/effect/complete를 한 transaction으로 제공하는 별도 execution profile은
|
|
실제 `TransactionPort` integration test를 통과한 경우만 `SAME_STORE_TRANSACTIONAL`을 광고한다.
|
|
|
|
### 24.6 owner-safe complete
|
|
|
|
```text
|
|
complete(scope, ownerToken, attempt, responseDigest, response, replayTtl)
|
|
```
|
|
|
|
program은:
|
|
|
|
1. record exists;
|
|
2. state is `EXECUTING`;
|
|
3. owner token matches;
|
|
4. attempt matches;
|
|
5. request fingerprint/policy compatible;
|
|
6. response size/version valid
|
|
|
|
를 모두 확인한 뒤 `COMPLETED`로 바꾼다.
|
|
|
|
stale owner는 new owner's record를 complete할 수 없다.
|
|
|
|
### 24.7 idempotent complete
|
|
|
|
같은 owner/attempt/response digest의 duplicate complete는 `ALREADY_COMPLETED_SAME_RESULT`로
|
|
성공 취급할 수 있다. 다른 digest는 conflict다.
|
|
|
|
이 규칙은 complete response loss 후 reconciliation을 돕는다.
|
|
|
|
### 24.8 release/fail
|
|
|
|
현재처럼 action이 RuntimeException을 던졌다고 무조건 record를 delete하지 않는다.
|
|
|
|
- business action이 시작되기 전 확정 실패: owner-safe release 가능;
|
|
- side effect가 없음을 application이 증명: `FAILED_RETRYABLE`;
|
|
- side effect가 발생했을 수 있음: `INDETERMINATE`, manual/domain reconciliation;
|
|
- stale owner: no-op/conflict.
|
|
|
|
`discard(scope)` API는 제거한다.
|
|
|
|
### 24.9 renew/takeover
|
|
|
|
long action은 owner-safe renew를 사용할 수 있다. renewal 실패/unknown이면 application은 더 이상
|
|
single owner라고 가정하지 않는다.
|
|
|
|
expired `CLAIMED`는 business action이 시작되지 않았으므로 attempt를 증가시키고 새 owner
|
|
token으로 takeover할 수 있다. old owner의 complete/release를 막는다.
|
|
|
|
expired `EXECUTING`은 effect가 commit되고 Redis complete만 빠졌을 수 있으므로 자동 takeover하지
|
|
않는다. claim은 `RECOVERY_REQUIRED`를 반환하고 record를 `ABANDONED/effect-unknown`으로
|
|
fail closed한다. domain receipt, DB unique operation row, downstream provider receipt 등으로:
|
|
|
|
- effect가 committed임을 확인하면 owner-safe reconciliation complete;
|
|
- no effect를 확인하면 explicit `ABANDONED -> CLAIMED` retry transition;
|
|
- 어느 쪽도 확인할 수 없으면 manual review/admission closed
|
|
|
|
를 선택한다. 단순 lease expiry는 safe retry 증거가 아니다.
|
|
|
|
`IdempotencyReconciliationPort`는 ordinary `IdempotencyStorePortV2`와 bean/type을 분리한다.
|
|
default template은 public reconciliation endpoint나 default bean을 만들지 않는다.
|
|
|
|
product가 명시적으로 추가한 authenticated operator/domain reconciliation use case만:
|
|
|
|
1. authorization과 separation-of-duty를 확인;
|
|
2. `IdempotencyEffectEvidenceVerifier`로 DB/domain/downstream source-of-truth를 조회;
|
|
3. verifier가 성공해 만든 opaque verified evidence를 bounded digest/reference로 변환;
|
|
4. actor digest, allowlisted reason, ticket/evidence reference, requestedAt을 durable audit store에
|
|
먼저 기록;
|
|
5. 별도 reconciliation port를 호출
|
|
|
|
한다. app-bootstrap은 이 authorized reconciler에만 reconciliation bean을 주입하고 ordinary
|
|
request executor에는 store port만 주입한다. ArchUnit/composition test가 web controller, 일반
|
|
use case, idempotency executor의 reconciliation port dependency를 금지한다.
|
|
|
|
Redis program은 receipt가 진실인지 판별하지 못하고 expected attempt/state revision, prior
|
|
evidence, audit operation ID의 CAS/replay만 보장한다. Redis audit TTL이 끝나도 필요한 actor/reason/
|
|
evidence trail이 사라지지 않도록 durable audit retention을 별도로 둔다. caller가 보낸 임의
|
|
evidence digest를 verifier 없이 port에 전달하는 경로는 금지한다.
|
|
|
|
### 24.10 cross-store crash gap
|
|
|
|
```text
|
|
Redis claim acquired
|
|
JDBC business transaction commits
|
|
process crashes
|
|
Redis complete not written
|
|
lease expires
|
|
retry executes business action again
|
|
```
|
|
|
|
Redis idempotency만으로 이 gap을 제거할 수 없다.
|
|
|
|
필요한 보완:
|
|
|
|
- DB unique constraint;
|
|
- domain operation ID;
|
|
- same-store inbox/idempotency;
|
|
- transactional outbox;
|
|
- downstream provider idempotency key;
|
|
- reconciliation.
|
|
|
|
### 24.11 provider profile
|
|
|
|
| Provider | 가능한 guarantee |
|
|
| --- | --- |
|
|
| Redis | low-latency request replay, declared durability/failover |
|
|
| JPA same source DB | same-store atomic claim/effect가 실제 한 transaction일 때 강화 가능 |
|
|
| external provider key | 해당 provider 범위의 dedup |
|
|
|
|
JPA와 Redis가 동일 port contract suite를 실행하더라도 descriptor의 guarantee는 같지 않다.
|
|
|
|
### 24.12 response storage
|
|
|
|
- maximum encoded bytes;
|
|
- sensitive field allowlist;
|
|
- no auth token/secret;
|
|
- content type/codec/version;
|
|
- status/header allowlist;
|
|
- digest;
|
|
- encryption requirement;
|
|
- replay TTL;
|
|
- legal/privacy retention.
|
|
|
|
large response 전체를 Redis에 넣지 않고 stable result reference를 저장할 수 있다. reference가
|
|
expired/deleted될 때 replay contract를 별도로 정의한다.
|
|
|
|
### 24.13 failure policy
|
|
|
|
idempotency store unavailable은 fail closed다. request를 그냥 실행하면 duplicate protection을
|
|
조용히 제거하게 된다.
|
|
|
|
읽기/claim timeout의 indeterminate 상태를 500 miss나 new claim으로 바꾸지 않는다. caller-facing
|
|
error mapping은 retriable 503/409 등 product contract에서 결정한다.
|
|
|
|
### 24.14 provider cutover
|
|
|
|
`jdbc -> redis`를 rolling deploy 중 단순 config flip하면 old pod와 new pod가 서로 다른 store에서
|
|
같은 scope를 claim할 수 있다.
|
|
|
|
선택:
|
|
|
|
1. traffic drain 후 atomic cutover;
|
|
2. dual-read/single-write migration coordinator;
|
|
3. versioned scope namespace와 client epoch;
|
|
4. maintenance window.
|
|
|
|
dual-write claim은 두 store 사이 atomic하지 않으므로 기본으로 사용하지 않는다.
|
|
|
|
### 24.15 existing JPA schema v2 migration
|
|
|
|
현재 `idempotency_record`의 v1 field는 scope, request hash, status, response payload/ref,
|
|
`created_at`, 단일 `expires_at` 중심이다. owner-safe v2를 위해 Flyway expand migration이 nullable
|
|
column을 먼저 추가한다.
|
|
|
|
```text
|
|
record_version
|
|
state_revision
|
|
owner_token
|
|
attempt
|
|
claim_operation_id
|
|
last_transition_operation_id
|
|
last_transition_kind
|
|
last_transition_result_digest
|
|
reconciliation_evidence_digest
|
|
processing_lease_until
|
|
replay_until
|
|
policy_revision
|
|
response_codec_id
|
|
response_codec_version
|
|
response_digest
|
|
failure_disposition
|
|
updated_at
|
|
```
|
|
|
|
migration sequence:
|
|
|
|
1. **Expand**: nullable column/index/check constraint를 추가하고 v1 code가 계속 읽을 수 있게 유지.
|
|
2. **Bridge reader deploy**: completed v1 row와 v2 row를 모두 읽되 아직 v2 claim을 쓰지 않음.
|
|
3. **Drain**: idempotent mutation admission을 잠시 닫고 maximum action/lease time을 기다린 뒤 live
|
|
`IN_FLIGHT` v1 row가 0임을 query/metric으로 증명.
|
|
4. **Backfill completed**:
|
|
`record_version=2`, `replay_until=expires_at`, codec/digest를 existing response에서 계산.
|
|
completed row에는 fake owner를 만들어내지 않음.
|
|
5. **Switch**: 모든 old writer가 drain된 뒤 v2 claim/owner-safe transition을 한 번에 활성화.
|
|
6. **Enforce**: new v2 `CLAIMED`/`EXECUTING` row의
|
|
owner/attempt/claim-operation/processing-lease non-null check와
|
|
owner-aware indexes 추가.
|
|
7. **Observe**: mismatch, stale-owner, dual reader, reaper 결과를 compatibility window 동안 관측.
|
|
8. **Contract**: old `expires_at`/status interpretation과 v1 API 제거는 다음 migration에서 수행.
|
|
|
|
rollback:
|
|
|
|
- v2 writer switch 전: schema-compatible old application rollback 가능;
|
|
- v2 writer switch 후: v1 writer로 rollback 금지, v2-compatible roll-forward/feature disable만 허용;
|
|
- emergency rollback이 필요하면 mutation admission을 닫고 v2 live owner를 drain/reconcile한 후 실행.
|
|
|
|
schema backfill과 `jdbc -> redis` provider cutover를 같은 release에서 수행하지 않는다.
|
|
|
|
## 25. Redis session profile
|
|
|
|
### 25.1 auth mode
|
|
|
|
security mode는 exclusive다.
|
|
|
|
```text
|
|
jwt
|
|
redis-session
|
|
```
|
|
|
|
JWT mode:
|
|
|
|
- stateless;
|
|
- no Redis session connection/repository/filter;
|
|
- bearer-only CSRF policy;
|
|
- Redis failure가 authentication에 영향 없음.
|
|
|
|
Redis session mode:
|
|
|
|
- server-side opaque session;
|
|
- multi-pod shared repository;
|
|
- session Redis role required;
|
|
- repository failure 시 fail closed/re-auth;
|
|
- CSRF/cookie/session fixation policy 필수.
|
|
|
|
### 25.2 ownership
|
|
|
|
Spring Session은 transport/security infrastructure다.
|
|
|
|
- Redis leaf: session role connection과 repository provider;
|
|
- inbound web: cookie/CSRF/security/session behavior;
|
|
- app-bootstrap: exclusive mode composition;
|
|
- application-core: Spring Session 타입 없음.
|
|
|
|
“모든 사용자 session revoke”가 business use case가 될 때만 framework-free
|
|
`SessionRevocationPort`를 application에 추가한다.
|
|
|
|
### 25.2.1 adapter-internal storage contract
|
|
|
|
Spring repository가 raw template command를 조합하지 않도록 Redis leaf 내부에만 다음 typed
|
|
storage contract를 둔다. application/shared/web에는 노출하지 않는다.
|
|
|
|
```java
|
|
interface VersionedRedisSessionStore {
|
|
SessionMutationAttempt newMutationAttempt();
|
|
SessionCreateOutcome create(SessionCreateCommand command);
|
|
SessionInspectionOutcome inspect(SessionInspectionCommand command);
|
|
SessionSaveOutcome saveIfLive(SessionSaveCommand command);
|
|
SessionTouchOutcome touchIfLive(SessionTouchCommand command);
|
|
SessionRevokeOutcome tombstoneAndDelete(SessionRevokeCommand command);
|
|
SessionRotateOutcome rotate(SessionRotateCommand command);
|
|
}
|
|
|
|
record SessionMutationAttempt(String operationId) {}
|
|
```
|
|
|
|
repository는 Redis send 전에 operation ID, expected/new revision, encoded payload digest를 command에
|
|
freeze하고 response-loss retry/inspect에 동일 값을 사용한다.
|
|
|
|
```text
|
|
Create:
|
|
CREATED | ALREADY_CREATED_SAME_OPERATION | EXISTS_CONFLICT |
|
|
TOMBSTONED | ABSOLUTE_EXPIRED | INDETERMINATE | UNAVAILABLE
|
|
Inspect:
|
|
LIVE_SAME_MUTATION | LIVE_OTHER | TOMBSTONED_SAME_OPERATION |
|
|
TOMBSTONED_OTHER | ABSENT | ABSOLUTE_EXPIRED | UNAVAILABLE
|
|
Save:
|
|
SAVED | ALREADY_SAVED_SAME_OPERATION | ABSENT | STALE_REVISION |
|
|
MUTATION_CONFLICT | TOMBSTONED | ABSOLUTE_EXPIRED |
|
|
INDETERMINATE | UNAVAILABLE
|
|
Touch:
|
|
TOUCHED | ALREADY_TOUCHED_SAME_OPERATION | TOUCH_NOT_DUE |
|
|
ABSENT | STALE_REVISION | MUTATION_CONFLICT | TOMBSTONED |
|
|
ABSOLUTE_EXPIRED | INDETERMINATE | UNAVAILABLE
|
|
Revoke:
|
|
REVOKED_AND_DELETED | TOMBSTONED_ABSENT |
|
|
ALREADY_REVOKED_SAME_OPERATION | STALE_REVISION |
|
|
OPERATION_CONFLICT | INDETERMINATE | UNAVAILABLE
|
|
Rotate:
|
|
ROTATED | ALREADY_ROTATED_SAME_OPERATION | OLD_ABSENT |
|
|
STALE_REVISION | OLD_TOMBSTONED | NEW_ID_CONFLICT |
|
|
ABSOLUTE_EXPIRED | INDETERMINATE | UNAVAILABLE
|
|
```
|
|
|
|
create/save response를 잃으면 같은 command 반복이 stored last-mutation operation ID와 payload
|
|
digest/revision을 비교해 `ALREADY_*_SAME_OPERATION`을 반환한다. inspect에서 same mutation이
|
|
확인되어도 payload digest와 resulting revision이 모두 일치해야 applied로 reconcile한다.
|
|
`LIVE_OTHER`, stale revision, 다른 tombstone은 blind overwrite/delete가 아니라 conflict다.
|
|
|
|
revoke response-loss는 tombstone operation ID를 inspect하고, rotate는 old tombstone과 new live
|
|
session의 operation/revision을 함께 확인한다. 둘 중 하나만 보이면 fail closed + reconciliation
|
|
대상이다. arbitrary retry에서 새 operation/session ID를 만들지 않는다.
|
|
|
|
### 25.3 repository choice
|
|
|
|
baseline은 index가 필요 없는 repository를 선택한다.
|
|
|
|
```text
|
|
RedisSessionRepository
|
|
```
|
|
|
|
이 baseline은 principal indexing, session-destroyed event, logout-all, concurrent-session-control이
|
|
필요 없는 profile로 제한한다.
|
|
|
|
principal lookup, concurrent-session control, logout-all index가 실제 필요할 때만:
|
|
|
|
```text
|
|
RedisIndexedSessionRepository
|
|
```
|
|
|
|
를 선택한다.
|
|
|
|
`RedisIndexedSessionRepository + Redis Cluster`는 stock 구현만으로 principal index cleanup,
|
|
logout-all, concurrent-session-control guarantee를 제공하지 않는다. 임의 한 node의 keyspace
|
|
event만 구독해 다른 shard event를 놓칠 수 있기 때문이다.
|
|
|
|
이 guarantee가 필요하면:
|
|
|
|
1. non-Cluster dedicated session deployment를 사용하거나;
|
|
2. 모든 primary event 구독, topology-change 재구독, durable reconciliation/reaper를 구현한 별도
|
|
indexed provider
|
|
|
|
중 하나를 선택한다. 단순 topology test 한 번으로 guarantee를 승격하지 않는다.
|
|
|
|
### 25.4 session value
|
|
|
|
session에는 최소 정보만 둔다.
|
|
|
|
- authentication/session metadata;
|
|
- CSRF token;
|
|
- bounded allowlisted attributes;
|
|
- created/last-access/absolute-expiry;
|
|
- security revision.
|
|
|
|
large business aggregate, arbitrary request object, persistence entity를 저장하지 않는다.
|
|
|
|
### 25.5 serializer
|
|
|
|
default JDK serialization을 사용하지 않는다.
|
|
|
|
- explicit serializer bean;
|
|
- allowlisted type set;
|
|
- schema/version envelope;
|
|
- no unrestricted polymorphic typing;
|
|
- N/N-1 rolling compatibility;
|
|
- maximum bytes/depth/collection elements;
|
|
- corrupt session은 invented auth가 아니라 invalidate + re-auth;
|
|
- security context library version upgrade test.
|
|
|
|
### 25.6 cookie
|
|
|
|
production setting:
|
|
|
|
- CSPRNG opaque session ID;
|
|
- `Secure`;
|
|
- `HttpOnly`;
|
|
- appropriate `SameSite`;
|
|
- bounded path/domain;
|
|
- no session ID in URL;
|
|
- proxy/TLS termination awareness;
|
|
- cookie name/environment collision 방지;
|
|
- rotation during privilege change.
|
|
|
|
cookie secret/value를 log하지 않는다.
|
|
|
|
### 25.7 CSRF
|
|
|
|
cookie-based authentication은 browser가 credential을 자동 전송하므로 CSRF protection이 필요하다.
|
|
|
|
- state-changing method protection;
|
|
- token storage/transport;
|
|
- CORS와 credential setting;
|
|
- logout CSRF;
|
|
- multi-tab behavior;
|
|
- error mapping
|
|
|
|
을 web contract test로 검증한다. JWT bearer-only mode의 기존 CSRF disable과 섞지 않는다.
|
|
|
|
### 25.8 fixation과 rotation
|
|
|
|
login, privilege elevation, sensitive re-authentication 후 session ID를 rotate한다. old ID는
|
|
더 이상 valid하지 않아야 한다.
|
|
|
|
rotation 중 attributes/TTL copy와 old key delete의 crash path를 test한다. two active IDs가 잠깐
|
|
허용되는지, old ID를 즉시 deny하는지 contract를 명시한다.
|
|
|
|
### 25.9 idle와 absolute expiry
|
|
|
|
두 경계를 분리한다.
|
|
|
|
```text
|
|
idle timeout
|
|
absolute lifetime
|
|
```
|
|
|
|
매 access가 idle TTL을 touch해도 absolute lifetime을 넘지 않는다. Redis TTL은 physical cleanup,
|
|
session metadata는 logical expiry를 확인한다.
|
|
|
|
stock Spring Session의 idle `maxInactiveInterval`만으로 absolute lifetime이나 logout tombstone을
|
|
제공한다고 가정하지 않는다. `VersionedRedisSessionRepository` decorator/custom repository가:
|
|
|
|
- `findById` 후 SecurityContext 사용 전에 `absoluteExpiresAt`과 revocation marker 확인;
|
|
- 위반 시 fail closed + delete;
|
|
- save/touch TTL을 `min(idleTimeout, absoluteExpiresAt - serverNow)`로 제한;
|
|
- session revision/tombstone CAS;
|
|
- serializer envelope
|
|
|
|
를 소유한다. 이 custom path가 비활성이라면 descriptor에서 absolute lifetime, atomic live touch,
|
|
no-resurrection guarantee를 제거한다.
|
|
|
|
R2 custom path는 §13.10의 `session-create`, `session-save-if-live`,
|
|
`session-touch-if-live`, `session-tombstone-and-delete`, `session-rotate` v1 manifest를 사용한다.
|
|
stock repository의 plain save/delete를 이 guarantee의 대체로 인정하지 않는다.
|
|
|
|
### 25.10 touch
|
|
|
|
매 request full session write는 write amplification을 만든다.
|
|
|
|
- changed attribute save;
|
|
- bounded touch interval;
|
|
- atomic live-check + TTL update;
|
|
- absolute expiry guard;
|
|
- touch failure policy;
|
|
- concurrent logout race
|
|
|
|
를 다룬다.
|
|
|
|
touch throttling은 configured idle timeout보다 충분히 짧아야 한다. 정확한 값은 SLO/traffic로
|
|
결정한다.
|
|
|
|
baseline Spring Session modes는:
|
|
|
|
```text
|
|
FlushMode.ON_SAVE
|
|
SaveMode.ON_SET_ATTRIBUTE
|
|
```
|
|
|
|
로 고정하고 write amplification/concurrent overwrite contract test를 둔다. 다른 mode는 별도
|
|
capability revision이다.
|
|
|
|
### 25.11 logout resurrection
|
|
|
|
다음 race를 test한다.
|
|
|
|
```text
|
|
Request A reads session
|
|
Request B logs out and deletes session
|
|
Request A finishes and saves stale session
|
|
```
|
|
|
|
logout은 tombstone/revision을 owner-safe atomic program으로 먼저 기록한 뒤 session을 삭제한다.
|
|
stale request의 save/touch는 tombstone/revision CAS에서 거절한다. tombstone TTL은 가능한 stale
|
|
request 최대 수명과 shutdown/drain budget보다 길고 bounded하다.
|
|
|
|
### 25.12 concurrent mutation
|
|
|
|
Redis session은 일반적으로 application-level serializable transaction을 제공하지 않는다.
|
|
동시 request의 attribute write는 last-write-wins/merge conflict가 날 수 있다.
|
|
|
|
- mutable business state를 session에 저장하지 않음;
|
|
- security-sensitive attribute에 revision/CAS;
|
|
- concurrent request contract test;
|
|
- lost update가 허용되는 attribute만 일반 save.
|
|
|
|
### 25.13 expiry event
|
|
|
|
keyspace expiry notification은 cleanup optimization이다. exact expiry trigger나 logout audit
|
|
source가 아니다. logical expiry는 read 시 검증하고 orphan index는 reaper가 reconcile한다.
|
|
|
|
### 25.14 failure
|
|
|
|
session Redis unavailable:
|
|
|
|
- existing auth를 invented/anonymous success로 바꾸지 않음;
|
|
- protected endpoint fail closed;
|
|
- user-facing re-auth/retry behavior;
|
|
- readiness degraded;
|
|
- bounded error storm logging;
|
|
- recovered connection에서 stale session resurrection 방지.
|
|
|
|
### 25.15 backup/restore
|
|
|
|
old session snapshot을 restore하면 revoked/expired session이 되살아날 수 있다.
|
|
|
|
- session key epoch;
|
|
- security revision;
|
|
- restore 후 global invalidation option;
|
|
- incident runbook;
|
|
- backup retention/privacy
|
|
|
|
가 필요하다. session backup이 항상 유용하다고 가정하지 않는다.
|
|
|
|
## 26. Pub/Sub, keyspace notification, Streams 경계
|
|
|
|
### 26.1 Pub/Sub 허용 범위
|
|
|
|
허용:
|
|
|
|
- L1 cache invalidation hint;
|
|
- live UI refresh hint;
|
|
- loss-tolerant internal signal.
|
|
|
|
금지:
|
|
|
|
- business event authoritative delivery;
|
|
- outbox replacement;
|
|
- payment/notification job;
|
|
- session revoke의 유일한 전달;
|
|
- exact cache consistency.
|
|
|
|
Pub/Sub은 at-most-once이며 disconnect 중 message를 replay하지 않는다.
|
|
|
|
### 26.2 keyspace notification
|
|
|
|
notification은:
|
|
|
|
- 기본 disabled일 수 있음;
|
|
- server CPU overhead;
|
|
- Cluster node-specific subscription;
|
|
- expiry 발생 시각 지연;
|
|
- disconnect loss
|
|
|
|
가 있다. session index cleanup이나 diagnostics 보조에만 사용한다.
|
|
|
|
### 26.3 client-side caching/tracking
|
|
|
|
Redis client-side tracking으로 L1 invalidation을 받을 수 있으나:
|
|
|
|
- client/library/topology compatibility;
|
|
- invalidation connection lifecycle;
|
|
- disconnect 시 local cache flush;
|
|
- failover/reconnect;
|
|
- redirect/Cluster;
|
|
- maximum tracked prefixes
|
|
|
|
를 검증한다. baseline off다.
|
|
|
|
### 26.4 Streams ownership
|
|
|
|
Redis Streams producer/consumer가 application messaging capability가 되면 existing cache leaf의
|
|
internal helper가 아니라 messaging semantic port 구현이어야 한다.
|
|
|
|
- outbound producer: outbound messaging provider;
|
|
- consumer group listener: future inbound messaging adapter;
|
|
- shared connection/runtime는 추출 가능한 internal library 또는 duplicated narrow config;
|
|
- cache Redis failure policy 재사용 금지.
|
|
|
|
### 26.5 Streams guarantee
|
|
|
|
consumer group은 pending entry와 ACK로 redelivery를 제공할 수 있지만:
|
|
|
|
- ack 전 crash -> duplicate;
|
|
- async replication/persistence loss;
|
|
- trim과 PEL interaction;
|
|
- poison message;
|
|
- consumer reclaim;
|
|
- single stream hot key;
|
|
- no DB+XADD transaction
|
|
|
|
이 남는다.
|
|
|
|
end-to-end는 at-least-once + idempotent consumer/inbox로 표현한다. producer dedup feature가 있어도
|
|
consumer side effect exactly-once를 뜻하지 않는다.
|
|
|
|
### 26.6 DB dual write
|
|
|
|
```text
|
|
DB commit
|
|
XADD
|
|
```
|
|
|
|
사이 crash gap은 Redis가 해결하지 않는다. DB가 source of truth이면 transactional outbox/CDC가
|
|
우선이다.
|
|
|
|
## 27. Topology design
|
|
|
|
### 27.1 exclusive topology
|
|
|
|
deployment 하나는 정확히 하나를 선택한다.
|
|
|
|
```text
|
|
standalone
|
|
sentinel
|
|
cluster
|
|
```
|
|
|
|
host list가 비어 있거나 두 topology field가 동시에 설정되면 startup failure다.
|
|
|
|
### 27.2 standalone
|
|
|
|
용도:
|
|
|
|
- local development;
|
|
- CI focused integration;
|
|
- product가 external HA를 제공하는 managed endpoint.
|
|
|
|
single process Redis를 production HA라고 부르지 않는다. managed proxy endpoint 뒤 topology는
|
|
operator attestation과 provider documentation으로 descriptor에 기록한다.
|
|
|
|
### 27.3 Sentinel
|
|
|
|
요구:
|
|
|
|
- master name;
|
|
- independent Sentinel endpoints;
|
|
- Redis data-node credentials와 Sentinel credentials 분리;
|
|
- TLS;
|
|
- failover timeout;
|
|
- client master rediscovery;
|
|
- old master partition test;
|
|
- topology event telemetry.
|
|
|
|
Sentinel은 failover를 자동화하지만 replication은 eventual/asynchronous다. partition된 old master에
|
|
acknowledged write가 합류 후 사라질 수 있다.
|
|
|
|
Sentinel 자체도 quorum/majority가 필요하다. 같은 node/failure zone에 Sentinel을 몰아놓고 HA라고
|
|
표현하지 않는다.
|
|
|
|
### 27.4 Cluster
|
|
|
|
Redis Cluster:
|
|
|
|
- 16,384 hash slots;
|
|
- database 0;
|
|
- node redirect (`MOVED`, `ASK`);
|
|
- multi-key/transaction/program same-slot requirement;
|
|
- topology refresh;
|
|
- uncovered slot/cluster state;
|
|
- primary/replica mapping
|
|
|
|
을 client가 이해해야 한다.
|
|
|
|
### 27.5 Cluster key validation
|
|
|
|
build/unit:
|
|
|
|
- key builder hash tag invariant;
|
|
- every program descriptor key count/slot rule.
|
|
|
|
integration:
|
|
|
|
- different slot multi-key가 expected failure;
|
|
- same tag 성공;
|
|
- reshard during traffic;
|
|
- `MOVED`/`ASK`;
|
|
- new/unknown node;
|
|
- failover 후 program availability.
|
|
|
|
### 27.6 topology refresh
|
|
|
|
Lettuce Cluster는 periodic + adaptive refresh를 명시적으로 설정한다.
|
|
|
|
관측:
|
|
|
|
- refresh count/reason;
|
|
- `MOVED`/`ASK`;
|
|
- persistent reconnect;
|
|
- unknown node;
|
|
- refresh failure;
|
|
- topology age.
|
|
|
|
managed/Kubernetes/NAT 환경에서 Redis가 advertise한 node address가 application pod에서 reachable한지
|
|
deployment conformance test로 확인한다.
|
|
|
|
### 27.7 read routing
|
|
|
|
기본:
|
|
|
|
- coordination/session/idempotency/lease/rate: primary only;
|
|
- cache: primary by default;
|
|
- stale-tolerant cache region만 replica read opt-in.
|
|
|
|
replica read는 latency/scale option이지 read-your-write를 보장하지 않는다. invalidation 직후 old
|
|
replica value를 읽을 수 있음을 region descriptor에 명시한다.
|
|
|
|
### 27.8 multi-region
|
|
|
|
active-active/multi-region Redis는 이 R2 baseline 밖이다.
|
|
|
|
검토해야 할 것:
|
|
|
|
- conflict resolution;
|
|
- local/global quota;
|
|
- session home region;
|
|
- idempotency scope;
|
|
- fencing token order;
|
|
- WAN partition;
|
|
- replication lag;
|
|
- data residency.
|
|
|
|
단일 region 설계를 DNS global endpoint로 바꿨다고 multi-region correctness가 되지 않는다.
|
|
|
|
## 28. Replication, persistence, failover guarantee
|
|
|
|
### 28.1 asynchronous replication
|
|
|
|
Redis replication은 기본적으로 asynchronous다. primary가 write ACK 후 replica 전파 전에 죽으면
|
|
promoted replica에 write가 없을 수 있다.
|
|
|
|
영향:
|
|
|
|
- completed idempotency record loss;
|
|
- session loss;
|
|
- rate token rollback;
|
|
- lease/fence record loss;
|
|
- duplicate lock holder;
|
|
- cache cold/stale.
|
|
|
|
각 capability descriptor가 이 결과를 명시한다.
|
|
|
|
### 28.2 `WAIT`
|
|
|
|
`WAIT`는 지정 replica가 write를 받은 acknowledgement를 기다려 loss probability를 줄일 수 있다.
|
|
그러나 strong consistency나 CP를 만들지 않는다.
|
|
|
|
- 대상 write/`FCALL`/`EVALSHA`와 정확히 같은 physical connection, 같은 primary에서 write 응답
|
|
직후 실행;
|
|
- Cluster에서는 target key slot을 소유한 primary connection을 명시적으로 고정;
|
|
- 별도 Spring template 호출 두 번으로 connection affinity를 추정하지 않고 dedicated
|
|
`RedisConnection`/native connection callback 사용;
|
|
- 특정 replica identity가 아니라 acknowledgement 개수만 요청;
|
|
- Lua/Function 또는 `MULTI/EXEC` 내부의 blocking acknowledgement로 사용 금지;
|
|
- production `timeout=0` 금지;
|
|
- 요구 수보다 작은 반환은 timeout/degraded;
|
|
- timeout에도 실제 replica 수가 일부 ACK했을 수 있음;
|
|
- failover selection/partition;
|
|
- write + WAIT 전체의 client response loss;
|
|
- throughput/latency cost.
|
|
|
|
strict profile의 optional acknowledgement 강화로만 표현한다.
|
|
|
|
### 28.3 `WAITAOF`
|
|
|
|
supported Redis version에서 `WAITAOF`는 local/replica AOF fsync acknowledgement를 강화할 수 있다.
|
|
역시 cross-store atomicity, failover selection, zero-loss를 보장하지 않는다.
|
|
|
|
`WAITAOF`도 대상 write와 같은 physical connection/primary에서 바로 호출하며 Function/Lua나
|
|
`MULTI/EXEC` 안에 넣지 않는다. local AOF acknowledgement를 요구하는 profile은 selected primary에
|
|
AOF가 실제 enabled라는 external attestation을 먼저 검증한다.
|
|
|
|
사용 여부는 capability descriptor에:
|
|
|
|
```text
|
|
replica acknowledgements
|
|
local AOF acknowledgements
|
|
timeout behavior
|
|
achieved acknowledgement count
|
|
```
|
|
|
|
로 기록한다.
|
|
|
|
write와 acknowledgement command 사이의 process/connection crash gap, acknowledgement response
|
|
loss는 여전히 `INDETERMINATE`다.
|
|
|
|
### 28.4 RDB
|
|
|
|
RDB snapshot:
|
|
|
|
- compact backup/startup;
|
|
- snapshot 간 write loss 가능;
|
|
- fork/COW memory와 latency;
|
|
- snapshot failure monitoring.
|
|
|
|
recomputable cache에 적합할 수 있다. session/idempotency RPO는 snapshot interval만 보고
|
|
“durable”이라고 표현하지 않는다.
|
|
|
|
### 28.5 AOF
|
|
|
|
AOF:
|
|
|
|
- fsync policy별 loss/latency trade-off;
|
|
- rewrite;
|
|
- disk space;
|
|
- corruption/recovery;
|
|
- fork/COW;
|
|
- write error.
|
|
|
|
`everysec`는 일반적으로 최근 구간 loss 가능성이 있다. exact maximum loss를 환경 검증 없이
|
|
단정하지 않는다.
|
|
|
|
### 28.6 no persistence
|
|
|
|
cache role은 no-persistence를 선택할 수 있다.
|
|
|
|
조건:
|
|
|
|
- 모든 값 재생성 가능;
|
|
- cold-start source capacity;
|
|
- startup warm strategy;
|
|
- cache loss alert severity;
|
|
- session/idempotency co-location 없음.
|
|
|
|
### 28.7 role별 baseline
|
|
|
|
| Role | Persistence | Replication | Effective claim |
|
|
| --- | --- | --- | --- |
|
|
| cache | optional | optional/replica read | recomputable, loss acceptable |
|
|
| coordination | explicit AOF/RPO | primary + replicas | low-latency state, loss still possible |
|
|
| session | product RPO에 맞는 AOF/HA | primary + replicas | session continuity best effort, re-auth recovery |
|
|
|
|
“AOF + replica = never lose”는 금지 문구다.
|
|
|
|
### 28.8 failover result
|
|
|
|
failover 중 client operation은:
|
|
|
|
```text
|
|
known not sent
|
|
sent and rejected
|
|
applied but response lost
|
|
applied on old primary then lost
|
|
replayed on new primary
|
|
```
|
|
|
|
중 하나일 수 있다. client error class 하나만으로 정확히 구분되지 않을 수 있으므로
|
|
capability-specific operation token과 reconciliation이 필요하다.
|
|
|
|
### 28.9 restore
|
|
|
|
backup restore 후:
|
|
|
|
- key schema/program version;
|
|
- expired logical record;
|
|
- session security epoch;
|
|
- idempotency replay horizon;
|
|
- fencing high watermark;
|
|
- rate policy revision;
|
|
- orphan namespace
|
|
|
|
를 reconcile한다. raw restore 성공이 application consistency 완료를 뜻하지 않는다.
|
|
|
|
## 29. Memory, eviction, big key, hot key
|
|
|
|
### 29.1 memory budget
|
|
|
|
Redis `maxmemory`를 container/node memory limit과 같게 두지 않는다.
|
|
|
|
별도 headroom이 필요한 항목:
|
|
|
|
- allocator fragmentation;
|
|
- replication backlog/buffer;
|
|
- AOF buffer/rewrite;
|
|
- fork copy-on-write;
|
|
- client input/output buffer;
|
|
- script/function memory;
|
|
- OS/page cache;
|
|
- TLS/client overhead.
|
|
|
|
일부 replication/AOF buffer는 eviction 비교에서 제외될 수 있다. `mem_not_counted_for_evict`를
|
|
포함해 effective headroom을 관측한다.
|
|
|
|
### 29.2 capacity input
|
|
|
|
region/capability마다 계산 input을 문서화한다.
|
|
|
|
```text
|
|
estimated key cardinality
|
|
average/p95/max key bytes
|
|
average/p95/max value bytes
|
|
Redis object/allocator overhead
|
|
TTL distribution
|
|
write/read rate
|
|
replication factor
|
|
growth rate
|
|
headroom factor
|
|
```
|
|
|
|
가짜 “몇 GB면 충분” 값을 skeleton에 넣지 않는다. 대신 startup/config는 maximum payload,
|
|
cardinality budget, queue bound처럼 안전에 필요한 upper bound를 요구한다.
|
|
|
|
### 29.3 eviction policy
|
|
|
|
cache role:
|
|
|
|
- `allkeys-lfu`: reusable hot value 유지에 일반적으로 적합;
|
|
- `allkeys-lru`: recency가 workload와 더 맞을 때;
|
|
- volatile policy: TTL 누락 key가 eviction 대상에서 빠질 위험을 이해한 경우만;
|
|
- `noeviction`: cache write OOM이 source load storm을 만들 수 있어 별도 설계 필요.
|
|
|
|
coordination/session role:
|
|
|
|
- `noeviction`;
|
|
- write OOM을 explicit failure로 받아들임;
|
|
- capacity alert와 scale/runbook;
|
|
- correctness record가 arbitrary eviction되지 않음.
|
|
|
|
policy 이름만 검사하지 않고 실제 role/data와 맞는지 검증한다.
|
|
|
|
### 29.4 OOM semantics
|
|
|
|
`noeviction`에서 memory limit을 넘는 write는 실패할 수 있다. 기존 read가 된다고 capability가
|
|
healthy한 것은 아니다.
|
|
|
|
- session create/touch OOM -> fail closed/readiness down;
|
|
- idempotency claim/complete OOM -> fail closed/indeterminate;
|
|
- rate-limit mutation OOM -> policy failure mode;
|
|
- lease acquire OOM -> no acquire;
|
|
- cache put OOM -> source response는 가능하지만 degraded.
|
|
|
|
Lua/Function도 low-memory에서 첫 write와 후속 write의 behavior를 real Redis로 test한다.
|
|
|
|
### 29.5 big key
|
|
|
|
big key는:
|
|
|
|
- network/event-loop latency;
|
|
- serialization allocation;
|
|
- replication;
|
|
- persistence;
|
|
- delete latency;
|
|
- failover/recovery
|
|
|
|
를 악화시킨다.
|
|
|
|
방어:
|
|
|
|
- encoded/decoded maximum;
|
|
- collection element maximum;
|
|
- bounded batch;
|
|
- compression upper bound;
|
|
- `UNLINK` maintenance;
|
|
- CI big-key negative test;
|
|
- operator `--keystats`/sampling.
|
|
|
|
### 29.6 hot key
|
|
|
|
hot key는 memory가 작아도 single shard CPU/network를 포화시킨다.
|
|
|
|
예:
|
|
|
|
- global rate counter;
|
|
- one popular cache object;
|
|
- one tenant hash tag;
|
|
- global session index;
|
|
- single Redis Stream.
|
|
|
|
metric에 raw key를 넣지 않고 bounded sampled key fingerprint/operator tool로 찾는다.
|
|
|
|
### 29.7 dangerous collection operations
|
|
|
|
regular runtime에서 금지:
|
|
|
|
```text
|
|
KEYS
|
|
unbounded HGETALL
|
|
unbounded SMEMBERS
|
|
unbounded LRANGE 0 -1
|
|
unbounded ZRANGE
|
|
unbounded XREAD without COUNT/block deadline
|
|
```
|
|
|
|
모든 collection operation은 maximum result count와 byte budget을 갖는다.
|
|
|
|
### 29.8 `SCAN`
|
|
|
|
`SCAN`도 free가 아니다.
|
|
|
|
- maintenance/admin path만;
|
|
- bounded COUNT hint;
|
|
- rate limit;
|
|
- cancellation/deadline;
|
|
- duplicate/missing observation 허용;
|
|
- mutation 중 exact snapshot 아님;
|
|
- Cluster node별 scan;
|
|
- report-only default;
|
|
- `UNLINK` batch와 backpressure.
|
|
|
|
request handler에서 wildcard invalidation에 사용하지 않는다.
|
|
|
|
### 29.9 fragmentation
|
|
|
|
`used_memory`, RSS, allocator fragmentation ratio를 함께 본다. fragmentation threshold를
|
|
universal constant로 고정하지 않고 version/workload baseline과 추세로 alert한다.
|
|
|
|
active defrag/allocator/server tuning은 deployment policy다. application이 바꾸지 않는다.
|
|
|
|
### 29.10 cache stampede under eviction
|
|
|
|
eviction이 급증하면 hit ratio 하락 -> source load -> cache refill -> eviction의 feedback loop가
|
|
생긴다.
|
|
|
|
관측/대응:
|
|
|
|
- evicted keys rate;
|
|
- hit/miss trend;
|
|
- source fallback concurrency;
|
|
- cache write rejected;
|
|
- hot/big key;
|
|
- admission;
|
|
- lower TTL가 아니라 memory/cardinality root cause;
|
|
- stale serve/load shedding.
|
|
|
|
## 30. Time와 expiration
|
|
|
|
### 30.1 clock ownership
|
|
|
|
capability별 clock:
|
|
|
|
| Capability | Enforcement clock |
|
|
| --- | --- |
|
|
| cache physical TTL | Redis |
|
|
| cache envelope freshness | Redis write time + application observation |
|
|
| rate limit | Redis `TIME` |
|
|
| lease validity | Redis TTL + client monotonic elapsed budget |
|
|
| idempotency processing/replay | Redis server time |
|
|
| session idle TTL | repository/Redis |
|
|
| session absolute lifetime | stored metadata + server/client validation |
|
|
|
|
pod wall-clock만으로 shared quota/lease를 계산하지 않는다.
|
|
|
|
### 30.2 absolute expiration
|
|
|
|
Redis는 expiry를 absolute wall-clock timestamp로 다룬다. server clock jump가 대량 즉시 expiry나
|
|
수명 연장을 만들 수 있다.
|
|
|
|
- NTP/clock monitoring;
|
|
- large clock step alert;
|
|
- rate/lease tests with time movement;
|
|
- `resetAt`는 server time에서 계산;
|
|
- client monotonic clock은 local deadline duration에 사용.
|
|
|
|
### 30.3 active/passive expiration
|
|
|
|
expired key는 access 시 passive하게, background sampling으로 active하게 제거된다. keyspace
|
|
notification 시각은 logical TTL boundary와 동일하지 않을 수 있다.
|
|
|
|
application은:
|
|
|
|
- read result/TTL로 logical expiry 확인;
|
|
- expiry event를 correctness trigger로 사용하지 않음;
|
|
- memory가 즉시 회수된다고 가정하지 않음.
|
|
|
|
### 30.4 TTL sentinel
|
|
|
|
expirable capability key에서:
|
|
|
|
```text
|
|
TTL = -1 -> corruption/policy violation
|
|
TTL = -2 -> absent
|
|
```
|
|
|
|
로 구분한다. `-1`을 immortal success로 그대로 두지 않고 capability별 repair/quarantine과 alert를
|
|
수행한다.
|
|
|
|
### 30.5 jitter와 legal/security expiry
|
|
|
|
cache freshness에는 positive/negative jitter를 적용할 수 있다. session absolute expiry,
|
|
credential revocation, compliance deadline에는 positive jitter로 수명을 늘리지 않는다.
|
|
|
|
### 30.6 long duration bound
|
|
|
|
TTL을 millisecond integer로 변환할 때:
|
|
|
|
- overflow;
|
|
- zero truncation;
|
|
- negative;
|
|
- provider maximum;
|
|
- policy maximum
|
|
|
|
을 validation한다. `Duration`을 `int` millisecond로 축소하지 않는다.
|
|
|
|
### 30.7 expiry race
|
|
|
|
read 직후 TTL이 만료될 수 있다. lease/session/idempotency는 “GET 성공했으므로 앞으로 TTL 동안
|
|
valid”라고 추정하지 않는다.
|
|
|
|
- owner-safe program에서 value와 TTL을 같이 확인;
|
|
- lease handle remaining validity;
|
|
- session logical expiry;
|
|
- idempotency state transition server-side time.
|
|
|
|
## 31. Client와 connection runtime
|
|
|
|
### 31.1 client selection
|
|
|
|
기본 provider는 Spring Data Redis + Lettuce다.
|
|
|
|
선정 이유:
|
|
|
|
- Spring Boot/Spring Data integration;
|
|
- standalone, Sentinel, Cluster;
|
|
- sync/async/reactive API;
|
|
- thread-safe shared native connection support;
|
|
- Spring Session과 connection factory integration;
|
|
- topology refresh와 reconnect telemetry 접근.
|
|
|
|
이는 application에 Spring Data abstraction을 노출한다는 뜻이 아니다.
|
|
|
|
### 31.2 alternatives
|
|
|
|
Jedis:
|
|
|
|
- blocking model과 explicit pool이 단순한 workload에 적합할 수 있음;
|
|
- 같은 semantic contract를 구현하는 별도 provider 후보;
|
|
- default로 두 client를 동시에 만들지 않음.
|
|
|
|
Redisson:
|
|
|
|
- high-level distributed object 제공;
|
|
- 별도 lifecycle/semantics/dependency cost;
|
|
- port contract 뒤의 optional provider.
|
|
|
|
### 31.3 dependency 직접 소유
|
|
|
|
Redis leaf는 broad starter에 기대지 않고 필요한 dependency를 직접 선언한다.
|
|
|
|
```text
|
|
spring-data-redis
|
|
lettuce-core
|
|
spring-boot-autoconfigure
|
|
micrometer-core (실제 instrumentation 소유 시)
|
|
```
|
|
|
|
Spring Session dependency ownership은 다음으로 고정한다.
|
|
|
|
| Module | Direct dependency | Responsibility |
|
|
| --- | --- | --- |
|
|
| `:adapter:outbound:cache-redis` | `spring-session-core`, `spring-session-data-redis` | repository, Redis storage, versioned serializer/decorator; Servlet type 금지 |
|
|
| `:adapter:inbound:web` | 사용 type이 있을 때 `spring-session-core` | cookie/CSRF/security integration; Redis leaf/package 의존 금지 |
|
|
| `:app-bootstrap` | composition type을 compile할 때 `spring-session-core` | exclusive auth-mode composition과 qualified bean wiring |
|
|
|
|
Spring Security/web dependency는 inbound web이 직접 소유한다. app-bootstrap과 inbound web이
|
|
Spring Session type을 compile하지 않으면 해당 `spring-session-core` dependency도 추가하지
|
|
않는다. 어떤 경우에도 Redis leaf의 transitive `implementation` leakage에 기대지 않는다.
|
|
|
|
Boot의 Redis/Session auto-configuration은 배제하거나 조건을 좁혀, unqualified global
|
|
`RedisConnectionFactory`, default session repository, classpath 기반 repository activation을
|
|
만들지 못하게 한다. role-qualified factory와 explicit repository configuration만 허용한다.
|
|
composition contract test는 cache/session factory 오주입, duplicate repository, JWT mode의
|
|
Session bean/connection 생성을 실패시킨다.
|
|
|
|
version은 Spring Boot BOM을 사용하되 lockfile로 고정한다.
|
|
|
|
### 31.4 one runtime per deployment
|
|
|
|
deployment ID마다:
|
|
|
|
- `RedisClient`;
|
|
- client resources/event loop;
|
|
- connection factory;
|
|
- topology settings;
|
|
- credential/TLS material;
|
|
- metrics scope;
|
|
- lifecycle
|
|
|
|
를 명확히 소유한다.
|
|
|
|
같은 endpoint/credential/topology인 role은 policy가 compatible할 때만 runtime을 공유한다.
|
|
|
|
### 31.5 connection types
|
|
|
|
일반 non-blocking command:
|
|
|
|
- thread-safe shared native connection 사용 가능;
|
|
- connection 수보다 in-flight/queue bound가 중요.
|
|
|
|
다음은 전용 connection/pool이 필요할 수 있다.
|
|
|
|
- blocking `XREAD`;
|
|
- Pub/Sub;
|
|
- transaction with connection affinity;
|
|
- long-running maintenance;
|
|
- stateful command mode.
|
|
|
|
blocking operation이 일반 cache/rate connection을 점유하지 않는다.
|
|
|
|
### 31.6 timeouts
|
|
|
|
최소 분리:
|
|
|
|
```text
|
|
DNS/connect timeout
|
|
TLS handshake timeout
|
|
pool acquire timeout
|
|
command timeout
|
|
capability overall deadline
|
|
blocking command timeout
|
|
shutdown timeout
|
|
```
|
|
|
|
하나의 global timeout으로 합치지 않는다.
|
|
|
|
관계:
|
|
|
|
```text
|
|
connect/command/acquire 각각 finite
|
|
capability overall deadline <= caller deadline
|
|
lease wait + work budget < caller deadline
|
|
blocking timeout < connection lifecycle timeout
|
|
```
|
|
|
|
exact default는 workload SLO로 조정하지만 production config는 finite upper bound를 요구한다.
|
|
|
|
### 31.7 request queue
|
|
|
|
reconnect 중 command를 무제한 buffer하면 Redis outage가 application heap outage로 바뀐다.
|
|
|
|
- disconnected request queue bounded;
|
|
- max in-flight bounded;
|
|
- capability별 bulkhead;
|
|
- queue full -> `OVERLOADED_BEFORE_SEND`;
|
|
- heap-based bound가 아니라 command/byte estimate 고려;
|
|
- queue depth/oldest age metric;
|
|
- required coordination은 빠르게 reject.
|
|
|
|
Cluster client의 theoretical queue upper bound는 Lettuce의 connection fan-out을 포함해 최소:
|
|
|
|
```text
|
|
requestQueueSize * ((clusterNodeCount * 2) + 1)
|
|
```
|
|
|
|
을 capacity input으로 사용하고, 실제 connection/runtime/bulkhead 수를 곱해 heap budget을
|
|
검증한다. 이 식은 memory 예약량의 충분조건이 아니며 command payload/response bytes도 더한다.
|
|
|
|
### 31.8 reconnect와 replay
|
|
|
|
Lettuce는 reconnect와 pending command replay behavior를 가진다. non-idempotent mutation이 다시
|
|
전송되면 duplicate effect가 날 수 있다.
|
|
|
|
baseline mutation runtime은 BOM이 선택한 Lettuce API의 semantics를 compatibility test로 확인한
|
|
뒤 connection 생성 전에 다음을 고정한다.
|
|
|
|
```text
|
|
autoReconnect = true
|
|
replayFilter(command -> true) // true == replay 대상에서 제외
|
|
disconnectedBehavior = REJECT_COMMANDS
|
|
requestQueueSize = finite bound
|
|
```
|
|
|
|
즉 driver-level pending replay는 모두 억제한다. retry-safe `GET` 등도 reconnect 후 capability
|
|
wrapper가 새 invocation으로 total deadline 안에서 재시도한다. 선택적 replay가 필요하면 raw command
|
|
type만 보지 않고 `FCALL/EVALSHA` program identity를 구분하거나 replay policy별 client/connection을
|
|
분리한다.
|
|
|
|
operation을 분류한다.
|
|
|
|
| Operation | retry/replay |
|
|
| --- | --- |
|
|
| GET/TTL | deadline 안의 bounded retry 가능 |
|
|
| idempotent delete desired-absent | bounded retry 가능 |
|
|
| set same value/version | 조건부 가능 |
|
|
| INCR/token consume | operation dedup 없으면 자동 replay 금지 |
|
|
| idempotency claim | same operation/owner token으로 reconcile |
|
|
| lease acquire | same owner token inspect |
|
|
| XADD/PUBLISH | messaging-specific dedup/at-least-once contract |
|
|
|
|
client global auto-replay만 믿지 않고 capability wrapper가 certainty를 분류한다.
|
|
|
|
startup descriptor와 reconnect integration test는 effective client options가 Lettuce default와
|
|
다름을 검증한다. 사용 중인 Lettuce version에 `replayFilter` semantics가 다르거나 없으면 release를
|
|
막고 별도 no-replay connection strategy를 구현한다.
|
|
|
|
### 31.9 cancellation
|
|
|
|
caller timeout으로 future를 cancel해도 server command가 이미 실행되었을 수 있다.
|
|
|
|
- local wait cancellation과 server execution 결과를 구분;
|
|
- mutation은 `INDETERMINATE`;
|
|
- connection을 무조건 close해 다른 multiplexed command를 해치지 않음;
|
|
- operation token reconciliation.
|
|
|
|
### 31.10 backoff
|
|
|
|
retry는:
|
|
|
|
- total deadline 안;
|
|
- exponential bounded backoff;
|
|
- jitter;
|
|
- maximum attempts;
|
|
- only retry-safe error/operation;
|
|
- circuit/topology state awareness
|
|
|
|
를 따른다. Redis timeout에 모든 request가 같은 즉시 retry를 하지 않는다.
|
|
|
|
### 31.11 circuit breaker
|
|
|
|
generic circuit breaker를 모든 Redis operation에 동일 적용하지 않는다.
|
|
|
|
- cache read: open 시 source fallback;
|
|
- strict rate/idempotency/session: open 시 fail closed;
|
|
- lease: no acquire;
|
|
- health probe가 breaker를 계속 열지 않게 분리;
|
|
- half-open traffic bounded;
|
|
- reconnect/topology failure와 중복 폭증 방지.
|
|
|
|
bulkhead/queue bound가 우선이고 breaker는 장애 전파 제어 수단이다.
|
|
|
|
### 31.12 event loop
|
|
|
|
Lettuce/Netty event loop에서:
|
|
|
|
- blocking DB call;
|
|
- heavy JSON encode/decode;
|
|
- compression;
|
|
- business logic;
|
|
- synchronous wait
|
|
|
|
를 실행하지 않는다. sync adapter도 underlying event loop와 caller thread 책임을 명확히 한다.
|
|
|
|
### 31.13 virtual threads
|
|
|
|
Java 21 virtual thread를 사용해도 Redis server/event-loop/connection queue capacity가 늘어나는 것은
|
|
아니다. 더 많은 concurrent caller가 queue를 포화시킬 수 있으므로 in-flight semaphore와 deadline은
|
|
그대로 필요하다.
|
|
|
|
### 31.14 DNS와 endpoint
|
|
|
|
- startup DNS validation;
|
|
- TTL/re-resolution;
|
|
- managed failover endpoint;
|
|
- IPv4/IPv6;
|
|
- certificate SAN;
|
|
- Cluster advertised address;
|
|
- Kubernetes service/NAT
|
|
|
|
를 topology profile별로 test한다. resolved IP를 영구 cache하는 custom code를 만들지 않는다.
|
|
|
|
### 31.15 client name
|
|
|
|
bounded client name:
|
|
|
|
```text
|
|
app + env + role + instance-short-id
|
|
```
|
|
|
|
raw hostname/user data를 넣지 않는다. operator가 `CLIENT LIST`/managed metrics에서 workload를
|
|
식별할 수 있게 한다.
|
|
|
|
## 32. Configuration design
|
|
|
|
### 32.1 top-level shape
|
|
|
|
상위 capability platform 설계와 같은 prefix를 사용한다. provider 정의는 연결 후보를
|
|
등록할 뿐 activation하지 않으며, `capabilities`의 provider/mode/binding 선택만 activation
|
|
SSOT다.
|
|
|
|
```yaml
|
|
ca-skeleton:
|
|
providers:
|
|
redis:
|
|
deployments:
|
|
cache-main:
|
|
topology: standalone
|
|
standalone:
|
|
endpoints:
|
|
- host: redis-cache.internal
|
|
port: 6379
|
|
database: 0
|
|
client-name: worklog-cache
|
|
authentication:
|
|
username: cache-runtime
|
|
password-ref: secret://redis/cache/password
|
|
tls:
|
|
enabled: true
|
|
verify-hostname: true
|
|
trust-bundle-ref: secret://redis/cache/ca
|
|
timeout:
|
|
connect: 1s
|
|
command: 250ms
|
|
acquire: 100ms
|
|
shutdown: 5s
|
|
queue:
|
|
max-in-flight: 1024
|
|
disconnected-behavior: reject
|
|
max-buffered-disconnected-requests: 0
|
|
topology-refresh:
|
|
periodic: 30s
|
|
adaptive: true
|
|
read:
|
|
preference: primary
|
|
|
|
coordination-main:
|
|
topology: sentinel
|
|
sentinel:
|
|
master-name: ca-coordination
|
|
endpoints:
|
|
- host: sentinel-a.internal
|
|
port: 26379
|
|
- host: sentinel-b.internal
|
|
port: 26379
|
|
- host: sentinel-c.internal
|
|
port: 26379
|
|
authentication:
|
|
username: coordination-sentinel-discovery
|
|
password-ref: secret://redis/coordination/sentinel-password
|
|
tls:
|
|
enabled: true
|
|
verify-hostname: true
|
|
trust-bundle-ref: secret://redis/coordination/sentinel-ca
|
|
authentication:
|
|
username: coordination-runtime
|
|
password-ref: secret://redis/coordination/password
|
|
tls:
|
|
enabled: true
|
|
verify-hostname: true
|
|
trust-bundle-ref: secret://redis/coordination/data-ca
|
|
timeout:
|
|
connect: 1s
|
|
command: 500ms
|
|
acquire: 100ms
|
|
shutdown: 5s
|
|
|
|
session-main:
|
|
topology: sentinel
|
|
sentinel:
|
|
master-name: ca-session
|
|
endpoints:
|
|
- host: session-sentinel-a.internal
|
|
port: 26379
|
|
- host: session-sentinel-b.internal
|
|
port: 26379
|
|
- host: session-sentinel-c.internal
|
|
port: 26379
|
|
authentication:
|
|
username: session-sentinel-discovery
|
|
password-ref: secret://redis/session/sentinel-password
|
|
tls:
|
|
enabled: true
|
|
verify-hostname: true
|
|
trust-bundle-ref: secret://redis/session/sentinel-ca
|
|
authentication:
|
|
username: session-runtime
|
|
password-ref: secret://redis/session/password
|
|
tls:
|
|
enabled: true
|
|
verify-hostname: true
|
|
trust-bundle-ref: secret://redis/session/data-ca
|
|
|
|
roles:
|
|
cache:
|
|
deployment: cache-main
|
|
required: false
|
|
expected-eviction: allkeys-lfu
|
|
coordination:
|
|
deployment: coordination-main
|
|
required: true
|
|
expected-eviction: noeviction
|
|
session:
|
|
deployment: session-main
|
|
required: true
|
|
expected-eviction: noeviction
|
|
|
|
programs:
|
|
mode: functions-provisioned
|
|
set-version: ca-redis-programs-v1
|
|
required-digest: sha256:...
|
|
|
|
key-digests:
|
|
default-profile: sensitive-scope
|
|
profiles:
|
|
sensitive-scope:
|
|
algorithm: hmac-sha-256
|
|
write-version: 2
|
|
readable-versions: [1, 2]
|
|
material-refs:
|
|
1: secret://redis/key-digest/hv1
|
|
2: secret://redis/key-digest/hv2
|
|
rotation-mode: dual-read-delete
|
|
maximum-read-probes: 2
|
|
coordination-scope:
|
|
algorithm: hmac-sha-256
|
|
write-version: 3
|
|
readable-versions: [3]
|
|
material-refs:
|
|
3: secret://redis/key-digest/coord-hv3
|
|
rotation-mode: cold-cutover
|
|
maximum-read-probes: 1
|
|
opaque-id:
|
|
algorithm: sha-256
|
|
write-version: 1
|
|
readable-versions: [1]
|
|
rotation-mode: fixed
|
|
maximum-read-probes: 1
|
|
```
|
|
|
|
role/deployment 항목의 존재만으로 client나 health bean을 만들지 않는다. 선택된 capability가
|
|
role을 참조할 때만 해당 runtime을 조립한다. 숫자는 starter example일 뿐 production SLO의
|
|
universal 정답이 아니다. typed validation과 environment-specific override가 필요하다. secret
|
|
값은 직접 YAML에 넣지 않는다.
|
|
|
|
### 32.2 topology sum type
|
|
|
|
Spring configuration binder가 sealed subtype을 자동 판별한다고 가정하지 않는다. binding model과
|
|
validated runtime model을 분리한다.
|
|
|
|
```text
|
|
@ConfigurationProperties("ca-skeleton.providers.redis")
|
|
RedisProviderProperties
|
|
Map<String, RedisDeploymentProperties> deployments
|
|
Map<RedisRole, RedisRoleBindingProperties> roles
|
|
RedisProgramSetProperties programs
|
|
RedisKeyDigestProperties keyDigests
|
|
|
|
RedisDeploymentProperties
|
|
topology: STANDALONE | SENTINEL | CLUSTER
|
|
standalone: StandaloneProperties?
|
|
sentinel: SentinelProperties? // master/endpoints + discovery authentication/TLS
|
|
cluster: ClusterProperties?
|
|
authentication/tls // selected data-node channel
|
|
|
|
RedisDeploymentSettingsFactory
|
|
-> StandaloneSettings | SentinelSettings | ClusterSettings
|
|
|
|
@ConfigurationProperties("ca-skeleton.capabilities")
|
|
CapabilitySelectionProperties
|
|
cache/rate-limit/idempotency/lock/security의 provider-neutral selection과 Redis role reference
|
|
```
|
|
|
|
`RedisDeploymentProperties`는 ordinary concrete `@ConfigurationProperties` record/class다.
|
|
factory가 discriminator와 exactly-one matching nested property를 검증하고 immutable runtime
|
|
sealed model을 만든다. non-selected nested property가 존재하거나 selected property가 빠지면
|
|
startup failure다.
|
|
|
|
binding test는 YAML/env -> properties -> factory -> exact runtime subtype 전 경로와 unknown/
|
|
contradictory field를 검증한다. custom Spring binder/converter는 이 단순 factory model로 표현할 수
|
|
없는 요구가 생길 때만 도입한다.
|
|
|
|
### 32.3 capability shape
|
|
|
|
```yaml
|
|
ca-skeleton:
|
|
capabilities:
|
|
cache:
|
|
bindings:
|
|
worklog-summary: redis
|
|
regions:
|
|
worklog-summary:
|
|
redis-role: cache
|
|
key-digest-profile: sensitive-scope
|
|
codec:
|
|
id: worklog-summary-json
|
|
write-version: 2
|
|
readable-versions: [1, 2]
|
|
maximum-payload: 256KiB
|
|
negative-ttl: 30s
|
|
soft-ttl: 8m
|
|
hard-ttl: 10m
|
|
ttl-jitter: 0.10
|
|
failure-mode: source-fallback
|
|
stampede: local-single-flight
|
|
|
|
rate-limit:
|
|
provider: redis
|
|
degraded-provider: local-emergency
|
|
redis-role: coordination
|
|
key-digest-profile: coordination-scope
|
|
local-emergency:
|
|
maximum-entries: 10000
|
|
entry-ttl: 2m
|
|
maximum-in-flight: 256
|
|
assumed-maximum-pods: 20
|
|
per-pod-share: 0.025
|
|
policies:
|
|
login:
|
|
revision: v3
|
|
algorithm: token-bucket
|
|
capacity: 10
|
|
refill-tokens: 10
|
|
refill-period: 1m
|
|
failure-mode: fail-closed
|
|
subject: [client-ip, route]
|
|
|
|
idempotency:
|
|
provider: redis
|
|
redis-role: coordination
|
|
key-digest-profile: coordination-scope
|
|
guarantee: request-replay
|
|
processing-lease: 30s
|
|
replay-ttl: 24h
|
|
maximum-response: 64KiB
|
|
|
|
lock:
|
|
bindings:
|
|
cache-refresh: redis
|
|
daily-export: jdbc
|
|
guarantees:
|
|
cache-refresh: cache-refresh-soft-lease
|
|
daily-export: efficiency-lease
|
|
redis-roles:
|
|
cache-refresh: cache
|
|
key-digest-profiles:
|
|
cache-refresh: sensitive-scope
|
|
|
|
security:
|
|
auth-mode: redis-session
|
|
session:
|
|
redis-role: session
|
|
key-digest-profile: opaque-id
|
|
repository: versioned-simple
|
|
idle-timeout: 30m
|
|
absolute-lifetime: 12h
|
|
touch-interval: 1m
|
|
serializer:
|
|
id: session-json
|
|
write-version: 3
|
|
readable-versions: [2, 3]
|
|
allowlisted-types:
|
|
- security-context-v1
|
|
- csrf-token-v1
|
|
maximum-payload: 64KiB
|
|
cookie:
|
|
name: WORKLOG_SESSION
|
|
secure: true
|
|
http-only: true
|
|
same-site: lax
|
|
path: /
|
|
domain: null # host-only; production fork가 필요한 경우에만 명시
|
|
csrf:
|
|
enabled: true
|
|
token-strategy: cookie-request-attribute
|
|
fixation:
|
|
strategy: migrate-session
|
|
rotate-on: [login, privilege-elevation, sensitive-reauth]
|
|
persistence:
|
|
flush-mode: on-save
|
|
save-mode: on-set-attribute
|
|
tombstone:
|
|
ttl: 5m
|
|
revision-cas: true
|
|
```
|
|
|
|
`bindings`, singleton `provider`, `dispatch-mode`, `auth-mode`가 각각 activation 축이다.
|
|
region/provider 내부에 별도 `enabled`를 두지 않는다. `disabled`를 선택하면 해당 capability가
|
|
비활성이다. 이 예시는 schema와 필수 보안 축을 고정하며 실제 product region/policy 값은 fork에서
|
|
정의한다.
|
|
|
|
### 32.4 validation
|
|
|
|
startup 전 deterministic validation:
|
|
|
|
- active capability에 provider/binding/mode 정확히 하나;
|
|
- rate-limit primary provider는 정확히 하나이며 degraded-provider는 `disabled` 또는 primary와 다른
|
|
provider 하나;
|
|
- referenced role/deployment 존재;
|
|
- topology exact one;
|
|
- endpoint non-empty/unique;
|
|
- TLS production requirement;
|
|
- Sentinel discovery channel과 discovered data-node channel의 named credential/trust material을
|
|
각각 표현하고 production에서 둘 다 검증;
|
|
- selected role의 data-node ACL username/password reference와 explicit trust bundle;
|
|
- secret reference 형식;
|
|
- timeout positive/order;
|
|
- queue/pool non-negative/positive relation (`reject` mode만 disconnected buffer 0 허용);
|
|
- Cluster database 0;
|
|
- role/read preference compatible;
|
|
- TTL relationships;
|
|
- algorithm parameter completeness;
|
|
- local-emergency maximum entries/TTL/in-flight positive bound와
|
|
`per-pod-share * assumed-maximum-pods <= 1`;
|
|
- key/codec/program version;
|
|
- key-digest algorithm/write/read versions, version별 material, rotation mode와 capability 호환성;
|
|
- provider guarantee가 requested guarantee 충족;
|
|
- incompatible role co-location;
|
|
- duplicate region/policy/provider ID;
|
|
- session serializer/cookie/CSRF/fixation/save/touch/tombstone 설정 완전성;
|
|
- `jwt`와 `redis-session` exclusivity 및 Redis Session에서 CSRF disable 거절;
|
|
- unknown setting fail closed where binder supports it.
|
|
|
|
### 32.5 runtime handshake
|
|
|
|
required capability activation:
|
|
|
|
1. connection/DNS/TLS/auth;
|
|
2. server role/topology;
|
|
3. supported Redis version;
|
|
4. Cluster coverage/database;
|
|
5. required command/program availability;
|
|
6. program digest/result schema;
|
|
7. read/write probe on dedicated ephemeral namespace;
|
|
8. role policy attestation;
|
|
9. serializer/key schema registry;
|
|
10. health registration.
|
|
|
|
probe key는 bounded TTL과 dedicated prefix를 사용하고 cleanup한다. user data namespace를 건드리지
|
|
않는다.
|
|
|
|
### 32.6 optional activation
|
|
|
|
optional cache가 unavailable이라고 전체 application startup을 반드시 막지는 않는다.
|
|
|
|
- config invalid/codec/program mismatch: startup fail;
|
|
- backend temporarily unavailable + declared optional: degraded startup 가능;
|
|
- session/idempotency/strict rate required: readiness/startup fail policy;
|
|
- descriptor에 actual state 표시.
|
|
|
|
misconfiguration과 external outage를 구분한다.
|
|
|
|
### 32.7 environment key registry
|
|
|
|
새 setting은:
|
|
|
|
- typed property;
|
|
- `application.yml` placeholder;
|
|
- env registry;
|
|
- `.env` example;
|
|
- binding/validation test;
|
|
- secret classification;
|
|
- documentation
|
|
|
|
을 한 change set에서 갱신한다.
|
|
|
|
현재 registry에만 있고 consumer가 없는 host/port/password/TTL key를 먼저 정리한다.
|
|
|
|
### 32.8 dynamic refresh
|
|
|
|
topology endpoint, credential rotation은 client lifecycle로 refresh할 수 있다. cache TTL/policy,
|
|
rate algorithm/revision, serializer/program version을 arbitrary live mutation하지 않는다.
|
|
|
|
policy 변경:
|
|
|
|
- new revision key;
|
|
- validate;
|
|
- shadow/canary;
|
|
- atomic registry switch;
|
|
- old state TTL drain.
|
|
|
|
Spring `@RefreshScope`로 connection/serializer가 중간 상태가 되게 하지 않는다.
|
|
|
|
## 33. Activation과 bootstrap
|
|
|
|
### 33.1 provider selection
|
|
|
|
선택은 typed ID로 한다.
|
|
|
|
```text
|
|
ca-skeleton.capabilities.cache.bindings.<region>=disabled|redis
|
|
ca-skeleton.capabilities.rate-limit.provider=disabled|local-emergency|redis
|
|
ca-skeleton.capabilities.rate-limit.degraded-provider=disabled|local-emergency
|
|
ca-skeleton.capabilities.idempotency.provider=disabled|jdbc|redis
|
|
ca-skeleton.capabilities.lock.bindings.<purpose>=disabled|local|jdbc|redis
|
|
ca-skeleton.capabilities.security.auth-mode=jwt|redis-session
|
|
```
|
|
|
|
`@Primary`나 classpath 우연으로 선택하지 않는다.
|
|
|
|
rate-limit의 exactly-one 규칙은 primary `provider`에 적용한다. `degraded-provider`는 별도 optional
|
|
축이고 primary와 같은 ID를 선택할 수 없다. selected primary가 Redis일 때만 Redis provider가,
|
|
selected degraded provider가 local-emergency일 때만 bounded local provider/composite가 생긴다.
|
|
|
|
### 33.2 disabled behavior
|
|
|
|
capability disabled:
|
|
|
|
- no provider bean;
|
|
- no client/connection;
|
|
- no scheduler/watchdog;
|
|
- no script load;
|
|
- no health dependency;
|
|
- no metric polling;
|
|
- direct use 시 typed `CapabilityDisabledException`.
|
|
|
|
empty optional cache router bean이 있다고 capability가 활성인 것은 아니다.
|
|
|
|
### 33.3 current flag migration
|
|
|
|
`APP_CACHE_REDIS_ENABLED`와 기존 `app.cache.*`/`app.redis.*`는 단계적으로 교체한다.
|
|
|
|
Phase:
|
|
|
|
1. old key를 canonical `ca-skeleton.capabilities.*`/`ca-skeleton.providers.redis.*`의 legacy
|
|
alias로만 읽고 deprecation log;
|
|
2. binding/provider/mode 없이 old enable만 true이면 activation을 거절;
|
|
3. old/new 값이 모순되면 precedence를 정하지 않고 startup failure;
|
|
4. canonical 값만 descriptor와 bean creation을 결정;
|
|
5. migration release 뒤 old key 제거;
|
|
6. env registry/public docs snapshot 갱신.
|
|
|
|
enable boolean 하나로 host/role/region/guarantee를 추측하지 않는다.
|
|
|
|
### 33.4 multi-instance safety
|
|
|
|
현재 bean-name list 검사를 capability descriptor validation으로 바꾼다.
|
|
|
|
예:
|
|
|
|
```text
|
|
required: rate-limit/login GLOBAL
|
|
actual: local fixed-window
|
|
-> startup failure
|
|
```
|
|
|
|
```text
|
|
required: cache-refresh CACHE_REFRESH_SOFT_LEASE
|
|
actual: redis lease, cache role
|
|
-> allowed
|
|
```
|
|
|
|
```text
|
|
required: inventory STRICT_COORDINATION
|
|
actual: redis efficiency lease
|
|
-> startup failure
|
|
```
|
|
|
|
plain `Object` bean으로 통과할 수 없어야 한다.
|
|
|
|
### 33.5 readiness composition
|
|
|
|
bootstrap이 enabled required capability의 health를 readiness group에 포함한다.
|
|
|
|
- optional cache backend down: ready + degraded detail;
|
|
- required session down: not ready;
|
|
- strict rate down: policy에 따라 not ready 또는 fail-closed serving;
|
|
- idempotency required mutation path down: not ready;
|
|
- unused role: health check 없음.
|
|
|
|
### 33.6 profile exclusivity
|
|
|
|
다음 contradiction을 거절한다.
|
|
|
|
- JWT stateless + Redis Session filter 동시;
|
|
- session mode + session role absent;
|
|
- session role bound to evictable cache;
|
|
- idempotency Redis + JPA provider both active;
|
|
- same lock purpose에 JDBC/Redis both active;
|
|
- rate provider Redis + local limiter가 silent primary;
|
|
- Functions mode + digest absent.
|
|
|
|
## 34. Security design
|
|
|
|
### 34.1 network
|
|
|
|
- public internet 직접 노출 금지;
|
|
- private endpoint/VPC/network policy;
|
|
- source security group 최소화;
|
|
- Redis client, replication, Cluster bus protection;
|
|
- management port 별도 통제;
|
|
- egress allowlist.
|
|
|
|
application-level password만으로 network exposure를 정당화하지 않는다.
|
|
|
|
### 34.2 TLS
|
|
|
|
production:
|
|
|
|
- TLS enabled;
|
|
- hostname verification enabled;
|
|
- trusted CA explicit;
|
|
- protocol/cipher policy;
|
|
- certificate expiry alert;
|
|
- SNI/managed endpoint test;
|
|
- Cluster/Sentinel 각 channel test;
|
|
- plaintext downgrade 금지.
|
|
|
|
`trust-all`이나 hostname verification off는 local-only이며 production startup에서 거절한다.
|
|
|
|
### 34.3 ACL identity
|
|
|
|
workload별 named user:
|
|
|
|
```text
|
|
cache-runtime
|
|
coordination-runtime
|
|
session-runtime
|
|
program-deployer
|
|
operator-readonly
|
|
```
|
|
|
|
default user는 production에서 disable한다.
|
|
|
|
### 34.4 least privilege
|
|
|
|
runtime user는 `reset -@all`에서 필요한 command/category/key/channel pattern만 부여한다.
|
|
|
|
금지 대상 예:
|
|
|
|
```text
|
|
CONFIG
|
|
ACL
|
|
MODULE
|
|
DEBUG
|
|
MONITOR
|
|
FLUSHALL/FLUSHDB
|
|
KEYS
|
|
MIGRATE
|
|
SHUTDOWN
|
|
FUNCTION LOAD/DELETE (runtime)
|
|
arbitrary EVAL (Functions profile)
|
|
```
|
|
|
|
ACL category가 새 Redis version에서 확장될 수 있으므로 allowlist와 negative integration test를
|
|
사용한다.
|
|
|
|
### 34.5 program deployment identity
|
|
|
|
Function provisioning account와 application runtime account를 분리한다.
|
|
|
|
- deployer: function library load/list/delete의 제한된 release workflow;
|
|
- runtime: `FCALL`과 data command;
|
|
- digest attestation;
|
|
- audit log;
|
|
- rollback artifact.
|
|
|
|
EVALSHA compatibility profile은 runtime script-load 권한의 위험을 capability card에 기록한다.
|
|
|
|
### 34.6 key pattern
|
|
|
|
ACL key pattern을 role/capability prefix에 제한한다. application key builder와 ACL pattern이
|
|
같은 versioned prefix registry에서 생성되도록 conformance test를 둔다.
|
|
|
|
hash tag/user input으로 prefix를 탈출할 수 없어야 한다.
|
|
|
|
### 34.7 secret source
|
|
|
|
password/private CA/key material은:
|
|
|
|
- secret reference;
|
|
- external secret manager/file mount;
|
|
- no source/YAML default;
|
|
- char/byte lifetime 최소화;
|
|
- structured log redaction;
|
|
- exception sanitization;
|
|
- rotation metadata.
|
|
|
|
현재 generic fail-open logger가 raw exception message를 기록하는 경로는 endpoint/credential
|
|
leak 가능성을 검토하고 classified sanitized field만 남기도록 바꾼다.
|
|
|
|
### 34.8 secret material provider contract
|
|
|
|
Redis leaf가 provider-specific SPI와 immutable value를 소유한다.
|
|
|
|
```java
|
|
public interface RedisCredentialMaterialProvider {
|
|
RedisCredentialResolution resolve(SecretReference reference);
|
|
RotationSubscription subscribe(
|
|
SecretReference reference,
|
|
RedisCredentialRotationListener listener);
|
|
}
|
|
|
|
public record VersionedRedisCredentialMaterial(
|
|
SecretVersion version,
|
|
Instant expiresAt,
|
|
DestroyableSecret username,
|
|
DestroyableSecret password,
|
|
DestroyableTrustMaterial trustMaterial) {}
|
|
```
|
|
|
|
`RedisCredentialResolution`은 `Resolved`, `TemporarilyUnavailable`, `Expired`,
|
|
`InvalidReference`, `PermissionDenied`를 구분한다. material은 version과 expiry를 가지며 사용 후
|
|
파기 가능한 byte/char representation으로 전달한다. secret value, reference 전체, provider
|
|
exception message를 metric/log에 남기지 않는다.
|
|
|
|
app-bootstrap은 환경에 맞는 Vault/file/Kubernetes/managed-secret 구현을 조립하거나 generic
|
|
secret capability를 Redis SPI에 bridge한다. Redis leaf가 bootstrap이나 특정 secret vendor에
|
|
역의존하지 않는다. listener는 새 version을 알릴 뿐 event thread에서 client를 직접 바꾸지 않고,
|
|
role runtime의 serialized rotation coordinator가 새 factory 검증, traffic switch, old connection
|
|
drain을 수행한다.
|
|
|
|
subscription loss, duplicate/out-of-order event, resolve timeout, expired material, partial role
|
|
rotation을 test한다. event만 믿지 않고 expiry 전 bounded periodic re-resolve를 둔다.
|
|
|
|
### 34.9 credential rotation
|
|
|
|
rotation protocol:
|
|
|
|
1. new credential/ACL 추가;
|
|
2. client dual-valid overlap;
|
|
3. new connection factory/session drain;
|
|
4. new credential connectivity/command test;
|
|
5. traffic switch;
|
|
6. old connections drain;
|
|
7. old credential revoke;
|
|
8. stale client alert.
|
|
|
|
한 global connection을 즉시 끊어 모든 role이 동시에 outage되지 않도록 role별 수행한다.
|
|
|
|
### 34.10 data at rest
|
|
|
|
AOF/RDB/backup에는 value가 평문으로 남을 수 있다.
|
|
|
|
- encrypted volume/managed KMS;
|
|
- backup encryption/access/retention;
|
|
- session/idempotency sensitive payload 최소화;
|
|
- application-level field encryption이 필요하면 별도 key rotation 설계;
|
|
- key names에도 PII 없음.
|
|
|
|
### 34.11 untrusted input
|
|
|
|
검증:
|
|
|
|
- key length;
|
|
- policy/region ID allowlist;
|
|
- cost upper bound;
|
|
- TTL upper bound;
|
|
- payload size/depth;
|
|
- collection count;
|
|
- script args;
|
|
- numeric overflow;
|
|
- Unicode normalization;
|
|
- compression ratio.
|
|
|
|
client가 Redis command name, key prefix, Lua source를 입력할 수 없다.
|
|
|
|
### 34.12 SSRF와 endpoint
|
|
|
|
Redis endpoint는 operator config에서만 온다. request/tenant가 host/port/database를 선택하지 않는다.
|
|
dynamic per-tenant Redis endpoint가 필요하면 별도 allowlisted tenancy control plane을 설계한다.
|
|
|
|
### 34.13 audit
|
|
|
|
audit 대상:
|
|
|
|
- provider/role binding change;
|
|
- function deploy/rollback;
|
|
- ACL/credential rotation;
|
|
- destructive maintenance;
|
|
- mass invalidation;
|
|
- idempotency manual reconciliation;
|
|
- fencing high-watermark repair;
|
|
- session global revoke.
|
|
|
|
audit에는 value/key/token/secret를 기록하지 않는다.
|
|
|
|
## 35. Health와 observability
|
|
|
|
### 35.1 health 의미
|
|
|
|
`PING` 성공만으로 다음을 보장하지 않는다.
|
|
|
|
- write 가능;
|
|
- correct primary;
|
|
- Cluster slot coverage;
|
|
- persistence 정상;
|
|
- noeviction headroom;
|
|
- required Function version;
|
|
- serializer/key compatibility;
|
|
- ACL command permission.
|
|
|
|
health는 capability와 role 관점으로 구성한다.
|
|
|
|
### 35.2 liveness
|
|
|
|
liveness는 Redis에 의존하지 않는다. Redis outage로 pod를 반복 재시작하면 connection storm과
|
|
failover를 악화시킨다.
|
|
|
|
### 35.3 readiness
|
|
|
|
required role:
|
|
|
|
- connection/auth/TLS;
|
|
- topology/primary;
|
|
- minimal read/write capability;
|
|
- program digest;
|
|
- recent success/error budget;
|
|
- queue saturation;
|
|
- role-specific requirement
|
|
|
|
을 본다.
|
|
|
|
optional cache는 readiness를 내리지 않을 수 있지만 `DEGRADED`를 표시한다.
|
|
|
|
### 35.4 capability metrics
|
|
|
|
공통:
|
|
|
|
```text
|
|
redis.capability.operations
|
|
redis.capability.duration
|
|
redis.capability.inflight
|
|
redis.capability.queue.depth
|
|
redis.capability.timeouts
|
|
redis.capability.indeterminate
|
|
```
|
|
|
|
bounded tags:
|
|
|
|
```text
|
|
deployment
|
|
role
|
|
capability
|
|
operation
|
|
outcome
|
|
topology
|
|
```
|
|
|
|
endpoint, key, tenant, user, session, owner token은 tag가 아니다.
|
|
|
|
### 35.5 cache metrics
|
|
|
|
```text
|
|
cache.lookup [hit, miss, negative, stale, unavailable, corrupt]
|
|
cache.write [stored, skipped, rejected, unavailable, indeterminate]
|
|
cache.invalidate
|
|
cache.source.load
|
|
cache.source.wait
|
|
cache.singleflight.join
|
|
cache.refresh.claim
|
|
cache.stale.age
|
|
cache.payload.bytes
|
|
```
|
|
|
|
region ID는 startup allowlist라 bounded tag로 허용할 수 있다.
|
|
|
|
### 35.6 rate metrics
|
|
|
|
```text
|
|
rate.decisions [allow, deny]
|
|
rate.enforcement [global, local-emergency, fail-open, fail-closed, shadow]
|
|
rate.algorithm
|
|
rate.indeterminate
|
|
rate.dedup.replay
|
|
rate.state.rejected
|
|
```
|
|
|
|
policy ID/revision은 bounded registry value다. subject는 tag/log에 넣지 않는다.
|
|
|
|
### 35.7 lease/fencing metrics
|
|
|
|
```text
|
|
lease.acquire [acquired, contended, unavailable, indeterminate]
|
|
lease.wait
|
|
lease.renew
|
|
lease.lost
|
|
lease.release [released, not-owner, indeterminate]
|
|
fence.issued
|
|
fence.rejected
|
|
fence.regression
|
|
```
|
|
|
|
resource digest도 metric tag로 쓰지 않는다.
|
|
|
|
### 35.8 idempotency metrics
|
|
|
|
```text
|
|
idempotency.claim [acquired, replay, in-progress, mismatch, unavailable]
|
|
idempotency.takeover
|
|
idempotency.renew
|
|
idempotency.complete
|
|
idempotency.owner-conflict
|
|
idempotency.indeterminate
|
|
idempotency.response.bytes
|
|
```
|
|
|
|
use-case/policy ID는 bounded registry일 때만 tag다.
|
|
|
|
### 35.9 session metrics
|
|
|
|
```text
|
|
session.load
|
|
session.save
|
|
session.touch
|
|
session.rotate
|
|
session.logout
|
|
session.corrupt
|
|
session.expired
|
|
session.reauth
|
|
session.repository.error
|
|
```
|
|
|
|
principal/session ID 없음.
|
|
|
|
### 35.10 client/topology metrics
|
|
|
|
```text
|
|
connect/reconnect
|
|
command timeout
|
|
queued/rejected command
|
|
pool acquire/saturation
|
|
MOVED/ASK
|
|
topology refresh/failure/age
|
|
sentinel failover
|
|
connection age
|
|
TLS/auth failure
|
|
NOSCRIPT
|
|
function digest mismatch
|
|
BUSY/slow program
|
|
```
|
|
|
|
### 35.11 server metrics
|
|
|
|
operator monitoring:
|
|
|
|
- `used_memory`, RSS, fragmentation;
|
|
- `mem_not_counted_for_evict`;
|
|
- `evicted_keys`, `expired_keys`;
|
|
- hit/miss;
|
|
- connected/blocked/rejected clients;
|
|
- replication role/link/lag/offset;
|
|
- AOF/RDB/rewrite/fork status;
|
|
- persistence error;
|
|
- commandstats/errorstats/latencystats;
|
|
- Cluster state/uncovered slots;
|
|
- slowlog/latency events;
|
|
- function/script memory/version.
|
|
|
|
application이 server INFO 전체를 high-cardinality metric으로 무분별하게 export하지 않는다.
|
|
|
|
### 35.12 tracing
|
|
|
|
span:
|
|
|
|
```text
|
|
redis capability operation
|
|
deployment/role
|
|
program name/version
|
|
outcome/certainty
|
|
duration
|
|
```
|
|
|
|
raw command argument/key/value를 기록하지 않는다. source cache load는 별도 child span으로 Redis
|
|
latency와 DB latency를 구분한다.
|
|
|
|
### 35.13 logs
|
|
|
|
structured event:
|
|
|
|
```text
|
|
event
|
|
capability
|
|
deployment/role
|
|
operation
|
|
outcome
|
|
errorCategory
|
|
certainty
|
|
correlationId
|
|
programVersion
|
|
```
|
|
|
|
raw exception message는 sanitize한다. repeated outage는 rate-limit/sampling하고 state transition은
|
|
반드시 남긴다.
|
|
|
|
### 35.14 alerts
|
|
|
|
최소 alert:
|
|
|
|
- required role unavailable;
|
|
- queue rejection/saturation;
|
|
- indeterminate mutation 증가;
|
|
- eviction on coordination/session;
|
|
- noeviction OOM;
|
|
- memory headroom;
|
|
- persistence failure;
|
|
- replication link/failover;
|
|
- Cluster uncovered slot;
|
|
- program digest drift/BUSY;
|
|
- session error/re-auth spike;
|
|
- idempotency owner conflict;
|
|
- lease lost/fence regression;
|
|
- rate local-emergency duration;
|
|
- cache miss/source load storm.
|
|
|
|
## 36. Lifecycle와 운영 제어
|
|
|
|
### 36.1 startup
|
|
|
|
순서:
|
|
|
|
1. typed config validation;
|
|
2. secret material resolution;
|
|
3. client resources;
|
|
4. topology/connect/auth;
|
|
5. program/schema capability;
|
|
6. role attestation;
|
|
7. provider binding;
|
|
8. health/readiness;
|
|
9. background refresh/watchdog/consumer.
|
|
|
|
background task를 connection validation 전에 시작하지 않는다.
|
|
|
|
### 36.2 graceful shutdown
|
|
|
|
순서:
|
|
|
|
1. readiness off/new traffic drain;
|
|
2. new cache refresh/rate background work 중단;
|
|
3. new lease/idempotency long operation 중단;
|
|
4. in-flight operation bounded wait;
|
|
5. owner-safe lease release best effort;
|
|
6. session save completion;
|
|
7. Pub/Sub/stream listener stop;
|
|
8. dedicated connection/pool close;
|
|
9. shared client resources close.
|
|
|
|
release response가 없다고 key를 blind delete하지 않는다.
|
|
|
|
### 36.3 deployment rollout
|
|
|
|
rollout compatibility 순서:
|
|
|
|
1. N reader가 N/N+1을 이해;
|
|
2. new program/function deploy;
|
|
3. digest 확인;
|
|
4. new application writer canary;
|
|
5. metrics/error 확인;
|
|
6. full rollout;
|
|
7. old payload/key/program TTL drain;
|
|
8. old reader/program 제거.
|
|
|
|
### 36.4 maintenance mode
|
|
|
|
destructive command는 application runtime에 없다.
|
|
|
|
operator tool/job:
|
|
|
|
- dry-run/report-only default;
|
|
- exact deployment/role/prefix;
|
|
- maximum keys/bytes;
|
|
- rate limit;
|
|
- approval/audit;
|
|
- resumable cursor;
|
|
- Cluster node coverage;
|
|
- `UNLINK` bounded batch;
|
|
- cancellation.
|
|
|
|
### 36.5 cache warmup
|
|
|
|
warmup은 optional:
|
|
|
|
- known bounded hot set;
|
|
- source load budget;
|
|
- randomized pacing;
|
|
- readiness와 분리;
|
|
- failure가 app startup을 무한 block하지 않음;
|
|
- no full DB/keyspace scan by default.
|
|
|
|
### 36.6 incident mode
|
|
|
|
capability별 safe degradation switch:
|
|
|
|
- cache: stale/source fallback budget;
|
|
- rate: fail closed/local emergency;
|
|
- session: re-auth/fail closed;
|
|
- idempotency: reject new mutation;
|
|
- lease: no new acquire;
|
|
- program mismatch: affected capability disable/fail.
|
|
|
|
global “ignore Redis errors” switch는 없다.
|
|
|
|
### 36.7 scaling
|
|
|
|
client pod scale-out 전에:
|
|
|
|
- Redis connection count;
|
|
- in-flight total;
|
|
- hot key;
|
|
- source fallback capacity;
|
|
- rate global key;
|
|
- session write amplification;
|
|
- topology refresh storm;
|
|
- credential/TLS handshake
|
|
|
|
를 계산한다. pod 수를 늘리면 Redis와 source가 자동 확장된다고 가정하지 않는다.
|
|
|
|
## 37. Test와 CI design
|
|
|
|
### 37.1 원칙
|
|
|
|
fake Redis는 application policy unit test에는 유용하지만 다음을 증명하지 못한다.
|
|
|
|
- command atomicity;
|
|
- TTL;
|
|
- wrong type;
|
|
- Lua/Function;
|
|
- script cache;
|
|
- Cluster slot;
|
|
- failover;
|
|
- replication/persistence loss;
|
|
- eviction/OOM;
|
|
- TLS/ACL;
|
|
- reconnect/replay.
|
|
|
|
R2 provider는 real Redis integration이 필수다.
|
|
|
|
### 37.2 application-core unit
|
|
|
|
framework/Redis 없이 hand-rolled fake port로:
|
|
|
|
- `CacheAsideExecutor` hit/miss/stale/unavailable;
|
|
- negative predicate;
|
|
- source failure/stale-if-error;
|
|
- `AuthoritativeAbsent`만 negative entry로 기록;
|
|
- `TransientFailure`/`PermanentFailure`/`Cancelled`는 negative entry로 기록하지 않음;
|
|
- stale은 `TransientFailure`에서만 policy에 따라 반환하고 permanent/unclassified failure에는 반환하지
|
|
않음;
|
|
- unclassified exception의 original cause/type를 보존하고 message를 cache/log/tag로 serialize하지
|
|
않음;
|
|
- local single-flight;
|
|
- source concurrency bound;
|
|
- idempotency claim outcome orchestration;
|
|
- owner lost/indeterminate;
|
|
- lease state/cancellation;
|
|
- business quota와 edge rate separation
|
|
|
|
을 검증한다.
|
|
|
|
### 37.3 shared edge contract
|
|
|
|
- request/decision validation;
|
|
- bounded policy/subject/evaluation ID;
|
|
- cost overflow;
|
|
- retry/reset semantics;
|
|
- enforcement/certainty enumeration;
|
|
- no Servlet/Redis dependency;
|
|
- serialization snapshot if wire/shared value로 노출될 때.
|
|
|
|
### 37.4 key builder unit/property
|
|
|
|
property test:
|
|
|
|
- same canonical input -> same key;
|
|
- different length-prefixed tuple -> collision 없음 within test corpus;
|
|
- raw PII substring 없음;
|
|
- maximum byte bound;
|
|
- invalid slug/braces 거절;
|
|
- HMAC version;
|
|
- rotation behavior;
|
|
- same resource atomic keys -> same slot;
|
|
- unrelated resource가 tenant-wide hot slot로 고정되지 않음;
|
|
- Cluster slot implementation과 real Redis `CLUSTER KEYSLOT` 일치.
|
|
|
|
### 37.5 codec contract
|
|
|
|
모든 codec:
|
|
|
|
- round-trip;
|
|
- deterministic form where required;
|
|
- N/N-1 read;
|
|
- future version reject;
|
|
- corrupt length/digest;
|
|
- maximum encoded/decoded;
|
|
- nested/decompression bomb;
|
|
- null/negative marker;
|
|
- forbidden polymorphic type;
|
|
- secret/PII fixture redaction;
|
|
- rolling writer/reader matrix.
|
|
|
|
JDK serialization marker나 native serialized payload가 fixture snapshot에 나타나면 실패한다.
|
|
|
|
### 37.6 program descriptor gate
|
|
|
|
build-time:
|
|
|
|
- every source/function has descriptor;
|
|
- descriptor checksum matches resource;
|
|
- unique name/version;
|
|
- explicit key count;
|
|
- result schema version;
|
|
- complexity/state bound non-empty;
|
|
- minimum Redis version;
|
|
- retry/certainty classification;
|
|
- no dynamic source concatenation;
|
|
- banned command/static pattern scan.
|
|
|
|
static scan은 semantic proof가 아니므로 real execution/concurrency test와 함께 사용한다.
|
|
|
|
### 37.7 standalone integration
|
|
|
|
Testcontainers real Redis에서:
|
|
|
|
- connection/settings;
|
|
- byte serializer;
|
|
- TTL on every expirable write;
|
|
- counter first-write/TTL concurrency에서 immortal key 없음;
|
|
- compare-delete/expire/set에서 stale owner/revision mutation 차단;
|
|
- hash/set/sorted-set/list/bitmap/HLL/geo primitive의 byte/cardinality/range/offset bound;
|
|
- unbounded collection API와 raw command/source facade가 public surface에 없음;
|
|
- cache hit/miss/negative/stale;
|
|
- conditional put/invalidate;
|
|
- namespace generation;
|
|
- program load/invoke/result;
|
|
- `NOSCRIPT` recovery;
|
|
- wrong-type/corrupt entry;
|
|
- rate algorithms;
|
|
- owner-safe lease;
|
|
- idempotency state machine;
|
|
- session repository;
|
|
- health/metrics.
|
|
|
|
container가 없으면 silently skip하지 않는 production-readiness task를 별도로 둔다.
|
|
|
|
### 37.8 cache concurrency
|
|
|
|
barrier-controlled tests:
|
|
|
|
- N simultaneous miss -> local loader once;
|
|
- multiple pod simulation -> bounded distributed refresh owner;
|
|
- lease expiry -> duplicate load 허용 but no corrupt put;
|
|
- invalidation during load -> old generation invisible;
|
|
- update revision vs stale put;
|
|
- Redis outage -> bounded source concurrency;
|
|
- eviction storm -> no unbounded thread/queue;
|
|
- corrupt entry -> no infinite reload loop;
|
|
- negative cache cardinality bound.
|
|
|
|
### 37.9 rate algorithm property
|
|
|
|
각 algorithm:
|
|
|
|
- exact/allowed approximation model과 reference implementation 비교;
|
|
- hundreds/thousands concurrent evaluation;
|
|
- window boundary;
|
|
- Redis server time movement;
|
|
- cost > 1;
|
|
- saturation/overflow;
|
|
- TTL cleanup;
|
|
- policy revision;
|
|
- evaluation dedup;
|
|
- response loss retry;
|
|
- Cluster same-slot;
|
|
- maximum state/member reject.
|
|
|
|
fixed window의 boundary burst는 bug로 무조건 실패시키지 않고 declared property로 검증한다.
|
|
sliding counter는 declared error bound를 검증한다.
|
|
|
|
### 37.10 lease/fencing concurrency
|
|
|
|
- only current owner releases;
|
|
- old owner after expiry cannot release/renew;
|
|
- acquire response loss + same token inspect;
|
|
- renew response loss;
|
|
- holder pause longer than TTL;
|
|
- lost callback/cancellation;
|
|
- failover duplicate-holder scenario;
|
|
- new fenced resource PENDING -> provision -> ACTIVE;
|
|
- ACTIVE counter missing은 request-path reinitialize 금지;
|
|
- restore는 durable `(epoch, highWatermark)` 이상에서만 reprovision;
|
|
- epoch/registration mismatch와 retired resource;
|
|
- protected resource rejects stale fencing token;
|
|
- counter regression availability behavior;
|
|
- `close()` duplicate call;
|
|
- shutdown during renew.
|
|
|
|
### 37.11 idempotency contract suite
|
|
|
|
JPA와 Redis provider 공통:
|
|
|
|
- first claim acquired;
|
|
- same fingerprint concurrent in progress;
|
|
- different fingerprint mismatch;
|
|
- completed replay;
|
|
- owner-safe renew/complete/release;
|
|
- stale owner blocked;
|
|
- expired `CLAIMED` takeover;
|
|
- expired `EXECUTING` -> `RECOVERY_REQUIRED`, no automatic re-execution;
|
|
- committed receipt reconciliation -> completed replay;
|
|
- authoritative no-effect evidence + expected revision -> reopened claim;
|
|
- stale/forged/conflicting reconciliation evidence CAS reject and audit;
|
|
- ordinary executor/controller에는 reconciliation port bean 주입 불가;
|
|
- verified evidence + authorized reconciler + durable audit 없이는 reconciliation 호출 불가;
|
|
- replay TTL separate;
|
|
- duplicate same complete idempotent;
|
|
- conflicting response digest rejected;
|
|
- oversize response;
|
|
- corrupt schema;
|
|
- unavailable/indeterminate.
|
|
|
|
provider-specific:
|
|
|
|
- Redis failover record loss/non-guarantee;
|
|
- JPA same-transaction claim/effect if advertised;
|
|
- JDBC/Redis cutover safety.
|
|
|
|
### 37.12 session integration
|
|
|
|
real Redis + multiple application contexts/pod simulation:
|
|
|
|
- create on pod A/read on B;
|
|
- create/save response loss -> same operation/digest replay or inspect reconciliation;
|
|
- conflicting mutation after lost response -> no blind overwrite;
|
|
- idle touch;
|
|
- absolute expiry;
|
|
- login ID rotation;
|
|
- old ID reject;
|
|
- logout/delete;
|
|
- concurrent stale save after logout;
|
|
- revoke/rotate partial observation -> fail closed reconciliation;
|
|
- concurrent attribute update;
|
|
- corrupt payload -> invalidate/re-auth;
|
|
- N/N-1 serializer;
|
|
- repository outage;
|
|
- noeviction OOM;
|
|
- failover;
|
|
- cookie/CSRF/security filter behavior;
|
|
- JWT mode has zero session Redis connection.
|
|
|
|
indexed repository는 Cluster/node-specific event와 orphan index cleanup을 별도 test한다.
|
|
|
|
### 37.13 Sentinel topology
|
|
|
|
최소 실제 topology:
|
|
|
|
- primary;
|
|
- replica 2개;
|
|
- 서로 다른 k3s node에 배치한 Sentinel 3개와 quorum 2.
|
|
|
|
test:
|
|
|
|
- client master discovery;
|
|
- primary kill;
|
|
- replica promotion;
|
|
- old primary partition/write;
|
|
- reconnect;
|
|
- in-flight mutation certainty;
|
|
- script/function availability;
|
|
- role/readiness event;
|
|
- credential/TLS.
|
|
|
|
단일 fake Sentinel endpoint로 HA를 증명하지 않는다.
|
|
|
|
#### 37.13.1 Sentinel discovery와 data runtime 분리
|
|
|
|
Sentinel discovery channel과 Redis data-node channel은 같은 Lettuce client/SSL context로
|
|
합치지 않는다. 각각 독립된 named material과 lifecycle을 갖는다.
|
|
|
|
| Channel | 책임 | 허용 material |
|
|
| --- | --- | --- |
|
|
| Sentinel discovery | master name 조회와 quorum 관측 | Sentinel ACL username/password reference, Sentinel CA/trust, discovery timeout |
|
|
| Redis data | capability command/program 실행 | data-node ACL username/password reference, data CA/trust, command/admission/drain timeout |
|
|
|
|
discovery는 다음 조건을 모두 만족할 때만 새 primary 후보를 반환한다.
|
|
|
|
- 구성된 Sentinel endpoint 최소 3개 중 2개 이상이 같은 master host/port를 보고한다;
|
|
- 응답한 Sentinel 수와 동의 수가 각각 bounded deadline 안에서 기록된다;
|
|
- master name이 exact configured name과 같다;
|
|
- 반환 endpoint가 loopback, wildcard, unspecified address가 아니고 allowlisted deployment
|
|
identity/member에 속한다;
|
|
- TLS hostname/SAN 검증을 통과한다;
|
|
- Sentinel credential 또는 trust를 data connection에, data material을 Sentinel connection에
|
|
재사용하지 않는다.
|
|
|
|
한 Sentinel의 응답, 최초 응답 또는 DNS 문자열 일치만으로 primary를 바꾸지 않는다. discovery
|
|
실패 detail에는 endpoint, username, secret reference/value, certificate subject를 남기지 않고
|
|
sanitized reason과 동의 수만 남긴다.
|
|
|
|
#### 37.13.2 bounded rediscovery와 runtime swap
|
|
|
|
정상 polling은 bounded single-flight로 실행하며, write/read command의 topology failure가
|
|
발생하면 같은 single-flight에 bounded immediate rediscovery를 요청한다. 새 primary가
|
|
qualification을 통과하면:
|
|
|
|
1. 새 data runtime을 생성한다;
|
|
2. version/program/semantic readiness를 검증한다;
|
|
3. 기존 `RedisRoleCommandRouter`에 한 번만 install한다;
|
|
4. 기존 runtime은 새 admission을 닫고 in-flight command를 bounded drain한다;
|
|
5. drain timeout 뒤에는 강제 close하되 완료되지 않은 mutation을 성공/미실행으로 추정하지 않는다.
|
|
|
|
failover 직전 또는 도중의 mutation은 자동 replay하지 않는다. transport가 실행 여부를 증명하지
|
|
못하면 capability가 `INDETERMINATE`를 반환하고, idempotency/session은 같은 operation token의
|
|
inspect/reconcile 또는 재인증 경로를 사용한다. read-only command도 semantic contract가 허용하는
|
|
경우에만 새 runtime에서 재시도한다.
|
|
|
|
`snapshot()`/readiness scrape는 정상 polling의 실행 엔진으로 사용하지 않는다. scrape나 command가
|
|
없는 동안에도 primary 변경을 발견해야 하므로, active Sentinel role이 하나 이상일 때만 registry가
|
|
다음 bounded poller를 소유한다.
|
|
|
|
- registry당 daemon worker 1개와 active Sentinel role당 fixed-delay task 1개만 만든다;
|
|
- 기본 polling period는 30초이고 typed setting은 5초 이상 5분 이하만 허용한다;
|
|
- scheduled poll과 command-failure trigger는 role별 같은 single-flight를 공유하며 한 role에
|
|
discovery/install 작업은 최대 1개만 실행하거나 대기한다;
|
|
- Standalone/Cluster만 선택되거나 Redis capability가 비활성이면 poller/thread/task를 0개 만든다;
|
|
- close는 새 trigger를 거절하고 scheduled task를 취소한 뒤 worker를 bounded shutdown하며,
|
|
close와 경합해 늦게 생성된 candidate는 install하지 않고 정확히 한 번 닫는다.
|
|
|
|
command failure signal은 route lease가 반환된 뒤 발행한다. connection/timeout/topology 계열의
|
|
`UNAVAILABLE`만 immediate rediscovery를 요청하고, overload, ACL denial, validation/size rejection은
|
|
요청하지 않는다. signal listener의 실패는 원래 command의 `NOT_APPLIED`/`INDETERMINATE` 판정을
|
|
절대 덮어쓰지 않는다.
|
|
|
|
정상 poll은 Sentinel discovery credential/CA만 사용해 endpoint를 조회한다. 현재 route와 같은
|
|
primary면 data credential/CA를 해석하거나 새 data connection을 열지 않는다. primary가 달라졌을
|
|
때만 이미 quorum-approved/allowlisted 된 exact endpoint로 data candidate를 열어 TOCTOU 성격의
|
|
이중 discovery를 피한다. route는 endpoint를 출력하지 않는 package-private identity와 monotonic
|
|
generation token을 가진다. candidate qualification 중 다른 rotation이 먼저 완료되면 stale
|
|
generation candidate를 닫고 install하지 않는다. 같은 identity도 candidate를 닫고 no-op 처리한다.
|
|
|
|
#### 37.13.3 replication 보장과 판정
|
|
|
|
Sentinel은 primary election을 제공하지만 asynchronous replication의 zero-data-loss를 보장하지
|
|
않는다. qualification 환경은 correctness role에 `min-replicas-to-write`와 bounded
|
|
`min-replicas-max-lag`를 설정하고, 중요한 mutation은 명시된 replica acknowledgement 정책을
|
|
사용한다. 이 설정도 strong consistency나 cross-store exactly-once 증거가 아니다.
|
|
|
|
failover 판정은 다음을 구분한다.
|
|
|
|
- 응답과 요구된 replica acknowledgement가 확인된 mutation: 새 primary에서 보존되어야 한다;
|
|
- response-only cut 또는 acknowledgement 결과를 확인할 수 없는 mutation:
|
|
`INDETERMINATE`, blind retry 금지;
|
|
- acknowledgement 전 명확한 connection/admission 실패: `NOT_APPLIED`가 wire evidence로
|
|
증명되는 경우에만 미실행으로 판정한다.
|
|
|
|
#### 37.13.4 Sentinel-first R2 qualification lab
|
|
|
|
이번 Phase 5의 첫 실행 slice는 기존 host k3s를 변경하지 않는 disposable Multipass lab이다.
|
|
|
|
```text
|
|
ca-redis-lab-server 2 CPU / 3 GiB / 12 GiB k3s server
|
|
ca-redis-lab-agent-1 2 CPU / 2.5 GiB / 12 GiB k3s agent
|
|
ca-redis-lab-agent-2 2 CPU / 2.5 GiB / 12 GiB k3s agent
|
|
pod CIDR 10.52.0.0/16
|
|
service CIDR 10.53.0.0/16
|
|
kube context ca-redis-lab
|
|
```
|
|
|
|
lab kubeconfig와 transient material/raw observation은 Gradle root의 ignored
|
|
`src/build/redis-lab` 아래에만 쓰며 사용자의 default kubeconfig에 merge하거나 덮어쓰지 않는다.
|
|
host 관측에는 default kubeconfig의 run-scoped copy와 시작 시점의 exact host context를
|
|
사용하지만, fingerprint/CIDR 관측이 끝난 즉시 성공/실패와 무관하게 copy를 제거한다. 모든
|
|
lab mutating command는 별도 lab kubeconfig와 `ca-redis-lab` context를 함께 요구한다.
|
|
VM 이름은 위 exact allowlist만 허용한다. launch 전에 exact name을 run-owned state에
|
|
`PENDING`으로 atomic 예약하고 성공 직후 `CREATED`로 승격한다. timeout, partial create,
|
|
state 승격 실패는 이 run이 예약한 exact name만 delete/purge한다. global `multipass purge`,
|
|
host `kubectl delete`, default-context write는 금지한다.
|
|
|
|
run-scoped rendered cloud-init은 secret이 아닌 exact `RUN_ID|VM_NAME` ownership marker를
|
|
instance에 기록한다. cleanup/down은 bounded marker read가 state owner와 name 일치를
|
|
증명할 때만 delete한다. launch timeout/error는 `RECONCILE` tombstone과 bounded late-create
|
|
poll로 처리한다. instance가 끝까지 없거나 marker가 unreadable/mismatch면 외부 same-name
|
|
instance를 추측해 삭제하지 않고 state를 유지한 채 fail-closed한다.
|
|
|
|
lifecycle 전체는 nonblocking exclusive lock과 run identity를 사용한다. direct `up`과
|
|
`run` 모두 첫 launch 전에 emergency cleanup을 활성화하며 signal/concurrent invocation이
|
|
다른 run의 state 또는 VM을 채택·삭제하지 못한다. `run -- <command>`에는 lifecycle lock file
|
|
descriptor를 상속하지 않는다. K3s는 mutable installer를 pipe로 실행하지 않고 exact release
|
|
URL/SHA-256을 repository에 pin한다. host download와 각 VM transfer 뒤 checksum/version을
|
|
다시 확인한 후에만 start한다.
|
|
기본 bounded external child도 lifecycle lock descriptor를 닫으며 lock acquisition만
|
|
명시적인 keep-lock 경로를 사용한다.
|
|
`run`의 inner `up` 성공과 user command 시작 사이에도 cleanup-required flag는 연속 유지되며,
|
|
signal handler가 ownership을 0으로 보는 handoff gap을 허용하지 않는다.
|
|
lab kubeconfig renderer는 one-cluster/context/user schema의 모든 identity-bearing key를
|
|
generic count하며 duplicate/extra server, context cluster/user, item/name,
|
|
current-context를 last-key-wins로 남기지 않고 fail-closed한다.
|
|
|
|
#### 37.13.4.1 lab lifecycle 완료 경계와 strict kubeconfig renderer
|
|
|
|
`Task 11.1A`는 하나의 리뷰 단위로 너무 많은 책임을 가졌으므로 다음 두 하위 작업으로 분리한다.
|
|
|
|
- `Task 11.1A-1`: VM 이름/소유권 marker, `PENDING|CREATED|RECONCILE` state, lock FD,
|
|
signal/handoff cleanup, host fingerprint와 bounded external command를 소유한다.
|
|
- `Task 11.1A-2`: pinned K3s admin kubeconfig의 strict validation과 lab 전용 rename/render만
|
|
소유한다.
|
|
|
|
`11.1A-1` 코드는 `11.1A-2` 동안 동결한다. `11.1A-2`가 독립 테스트와 독립 리뷰를 통과하기
|
|
전에는 부모 `11.1A`를 완료로 표시하지 않으며 VM 생성도 허용하지 않는다.
|
|
|
|
`11.1A-2`는 범용 YAML parser가 아니다. 입력은 pinned K3s가 생성하는 admin kubeconfig의
|
|
canonical block-style 문서 하나로 제한한다. 별도 tracked
|
|
`infra/redis-lab/lib/render-kubeconfig.awk`가 line/indentation/state allowlist를 적용하며,
|
|
identity-bearing key를 찾는 denylist나 발견된 mutation별 정규식 패치를 사용하지 않는다.
|
|
|
|
허용 grammar는 다음을 모두 만족해야 한다.
|
|
|
|
- top-level `apiVersion`, `clusters`, `contexts`, `current-context`, `kind`, `preferences`,
|
|
`users`는 canonical 순서와 exact spelling/indentation으로 한 번만 존재한다;
|
|
- cluster/context/user list는 각각 한 항목만 가지며 identity는 모두 exact `default`다;
|
|
- cluster는 exact loopback `server: https://127.0.0.1:6443`와 하나의
|
|
`certificate-authority-data` scalar만 가진다;
|
|
- context는 exact `cluster: default`, `user: default`와 optional single `namespace` scalar만
|
|
가진다;
|
|
- user는 하나의 `client-certificate-data`와 `client-key-data` scalar만 가진다;
|
|
- `preferences: {}`만 유일한 flow collection 예외다. 그 밖의 `{}`, `[]`, quoted/tagged/
|
|
explicit key, anchor, alias, merge key, tab, CRLF, YAML document marker, unknown key,
|
|
duplicate/reordered identity, trailing content는 fail-closed한다;
|
|
- source `server`, cluster/context/user name과 current-context만 변환한다. CA/client material,
|
|
namespace와 그 밖의 허용 scalar는 byte-preserving pass-through다;
|
|
- renderer source 자체와 destination의 canonical parent/symlink/permission 계약을 lifecycle
|
|
static validation에 포함한다. validation 또는 render 실패 시 destination을 제거하고
|
|
constant sanitized failure만 출력한다.
|
|
|
|
정상 fixture는 pinned K3s admin kubeconfig의 certificate-data shape를 사용한다. negative
|
|
mutation은 duplicate/extra identity뿐 아니라 canonical item 아래의 sibling
|
|
`cluster : {...}`, `context : {...}`, whitespace-before-colon, flow collection, quoted/tagged/
|
|
anchor/alias/merge, unknown/reordered/missing key를 포함한다. 모든 실패는 lab `kubectl` 전에
|
|
발생하고 현재 invocation이 marker로 증명한 VM만 cleanup하며 prior
|
|
`CREATED|RECONCILE` state는 byte-for-byte 보존한다.
|
|
|
|
tracked `infra/redis-lab`에는 lifecycle script, cloud-init template, Redis/Sentinel config
|
|
template, Kubernetes manifest와 secret 없는 contract test만 둔다. 실행 시 생성하는 k3s token,
|
|
ACL password, data/Sentinel/untrusted CA와 private key, rendered Secret/config, raw observation은
|
|
`umask 077`인 transient directory에만 둔다. `redis-cli --pass`, tracked PEM/Secret data,
|
|
`hostPath`/`hostNetwork`/privileged/NodePort/LoadBalancer는 사용하지 않는다.
|
|
|
|
host isolation은 preflight/postflight의 canonical projection을 비교한다. default kubeconfig
|
|
digest, current context/API, sorted node/providerID/podCIDR, controller replica, Service NodePort,
|
|
host interface/route CIDR와 Multipass inventory가 대상이다. host service CIDR은 현재 할당된
|
|
ClusterIP만 보고 추측하지 않고, 명시적으로 검증한 input 또는 신뢰할 수 있는 host 설정에서
|
|
읽는다. 외부 명령과 exact 3-node Ready 대기는 bounded다. 불일치 시 qualification을
|
|
실패시키되 script가 host 상태를 추측해 되돌리려고 mutate하지 않는다.
|
|
|
|
workload는 Redis primary 1 + replica 2, Sentinel 3/quorum 2를 서로 다른 node에 배치한다.
|
|
data와 Sentinel은 stable ordinal/headless DNS가 필요한 별도 StatefulSet이며
|
|
`kubernetes.io/hostname` required anti-affinity와 `maxSkew=1/DoNotSchedule` topology spread를
|
|
사용하고 `podManagementPolicy: Parallel`을 명시한다. data는 PVC와 AOF
|
|
`appendfsync everysec`를 사용한다. Sentinel config는 discovery/failover 시 rewrite되므로
|
|
bootstrap 원본을 pod별 writable PVC config로 최초 1회 atomic init-copy하되 restart 때 이미
|
|
존재하는 rewritten config를 덮어쓰지 않는다. 비어 있거나 손상된 기존 config도 자동으로
|
|
덮지 않고 startup을 실패시켜 증거를 보존한다.
|
|
|
|
data/Sentinel plaintext port는 0이며 TLS port만 연다. `tls-replication yes`, hostname
|
|
resolution/announcement와 stable DNS SAN을 사용한다. data plane과 Sentinel plane의 CA/leaf
|
|
material은 분리하며 peer 연결에 필요한 root만 explicit trust bundle에 포함한다. ACL은
|
|
application data, replica, Sentinel-to-data, Sentinel peer, application Sentinel discovery
|
|
identity로 나눈다. Redis data ACL과 Sentinel ACL은 별도 template/projection이며 plane
|
|
identity를 서로 노출하지 않는다. default user는 off이며 application/data/discovery
|
|
identity에는 `+@all`, `allkeys`, `allchannels`를 주지 않는다. replica는
|
|
`+psync +replconf +ping`, Sentinel-to-data identity는 Sentinel control에 필요한 최소
|
|
command/channel set만 가진다.
|
|
|
|
exec probe를 사용하고 default-deny NetworkPolicy 뒤 data 6379, Sentinel 26379, kube-dns,
|
|
exact qualification/application pod selector만 허용한다. data/Sentinel PDB는 각각
|
|
`minAvailable: 2`이며 non-root, read-only root filesystem, privilege-escalation false,
|
|
capability drop ALL, seccomp RuntimeDefault, requests/limits를 요구한다. `hostPath`,
|
|
host namespaces, privileged, NodePort/LoadBalancer와 tracked Secret/PEM은 금지한다.
|
|
|
|
정적 lifecycle contract와 manifest/security contract는 VM 없이 blocking check에서 검증하고,
|
|
한 필드씩 제거/변조하는 mutation-negative fixture로 실제 방어력을 확인한다. 이 정적 통과는
|
|
TLS handshake, ACL authorization, CNI enforcement, scheduling/failover의 실행 증거가 아니다.
|
|
shell contract는 별도 fixture repository만 사용하며 actual `src/build/redis-lab` state를
|
|
byte-for-byte 보존한다. fake PATH는 explicit safe wrapper 외 모든 명령을 fail-closed한다.
|
|
live lab에서는 TLS/ACL negative test, `SENTINEL CKQUORUM`, writable config rewrite/restart,
|
|
exact 3 Ready placement, PDB/NetworkPolicy enforcement와 image ID/digest를 별도로 검증한다.
|
|
Redis image는 `src/gradle/redis-test-images.properties`의 `redis.minimum.image` exact
|
|
tag+digest를 사용한다.
|
|
|
|
ordinal bootstrap은 최초 `redis-data-0` primary와 두 replica만 정적으로 증명한다.
|
|
failover 동안 죽어 있던 old primary가 재합류할 때 readiness가 stale direct write를 허용하지
|
|
않고 새 primary의 replica로 수렴하는지는 live gate다. PDB 선언은 voluntary eviction
|
|
제약일 뿐 node/AZ failure 증거가 아니다.
|
|
|
|
k3s control-plane HA, physical host/AZ failure, Redis Cluster는 이 lab의 증거가 아니다.
|
|
hosted GitHub Actions에서는 Multipass를 설치하거나 실행하지 않는다. 실제 lab qualification은
|
|
trusted dedicated runner 또는 local explicit execution에서만 허용한다. 외부 PR 코드를
|
|
self-hosted lab에서 실행하지 않는다.
|
|
|
|
초기 test budget은 운영 SLA가 아니라 bounded regression limit이다.
|
|
|
|
- Sentinel election: 60초 이내;
|
|
- client rediscovery와 runtime swap: election 뒤 추가 30초 이내;
|
|
- required semantic readiness 복구: fault injection 뒤 총 90초 이내.
|
|
|
|
실제 측정값을 evidence timeline에 기록하며 limit만 기록한 문서는 증거가 아니다.
|
|
|
|
#### 37.13.5 Sentinel-first capability acceptance
|
|
|
|
이 slice는 correctness-sensitive cross-pod state를 우선 검증한다.
|
|
|
|
- edge rate limit: failover 전 quota state가 조용히 reset되지 않고 evaluation replay가 일관된다;
|
|
- request-replay idempotency: claim/start/renew/complete와 terminal replay가 owner-safe하며
|
|
불확실 mutation은 중복 실행하지 않는다;
|
|
- Redis session: create/read/touch/rotate/revoke가 서로 다른 application pod에서 보이고,
|
|
failover 뒤 confirmed state가 유지되며 stale session이 부활하지 않는다;
|
|
- cache refresh soft lease와 optional cache는 공통 runtime 회귀를 확인하되 이 slice만으로
|
|
Cluster scaling 또는 distributed L1 invalidation R2를 주장하지 않는다.
|
|
|
|
fault 순서는 baseline qualification 뒤 current primary pod를 kill하고 readiness unavailable,
|
|
Sentinel quorum election, client rediscovery, runtime swap/drain, semantic readiness recovery를
|
|
실제 timestamp로 수집한다. old primary는 replica로 재합류해야 하고, recovery 뒤 모든 actor가
|
|
같은 runtime generation을 관측해야 한다.
|
|
|
|
evidence bundle은 실제 실행 image digest/image ID, config/program digest, fault/election/recovery
|
|
timeline, capability별 outcome/certainty, sanitized Kubernetes/Sentinel observation, lab teardown
|
|
결과를 포함한다. manifest의 `NOT_CAPTURED`를 문자열로 바꾸는 것만으로 증거를 만들 수 없다.
|
|
|
|
### 37.14 Cluster topology
|
|
|
|
최소 multi-primary Cluster와 replica에서:
|
|
|
|
- slot coverage;
|
|
- same/cross-slot;
|
|
- `MOVED`/`ASK`;
|
|
- reshard;
|
|
- primary failover;
|
|
- topology refresh;
|
|
- advertised address;
|
|
- program on every primary;
|
|
- Pub/Sub/tracking profile;
|
|
- DB 0 validation;
|
|
- bounded redirects.
|
|
|
|
### 37.15 TLS/ACL
|
|
|
|
- trusted CA succeeds;
|
|
- untrusted CA fails;
|
|
- hostname mismatch fails;
|
|
- plaintext profile rejected in prod;
|
|
- wrong/rotated credential;
|
|
- old/new overlap;
|
|
- cache user cannot touch session prefix;
|
|
- runtime cannot `CONFIG`, `KEYS`, `FLUSH*`, function deploy;
|
|
- program deployer cannot read application values beyond required;
|
|
- `ACL DRYRUN` or equivalent conformance;
|
|
- exception/log secret redaction.
|
|
|
|
### 37.16 memory/eviction
|
|
|
|
real server config:
|
|
|
|
- cache `allkeys-*` eviction;
|
|
- coordination/session `noeviction`;
|
|
- OOM write outcome;
|
|
- existing read behavior;
|
|
- Lua under memory pressure;
|
|
- big key rejection;
|
|
- eviction metric;
|
|
- cache and correctness deployment isolation;
|
|
- headroom alert inputs;
|
|
- `UNLINK` bounded cleanup.
|
|
|
|
### 37.17 persistence/restart
|
|
|
|
profiles:
|
|
|
|
- no persistence;
|
|
- RDB;
|
|
- AOF configured mode.
|
|
|
|
test:
|
|
|
|
- graceful restart;
|
|
- kill/power-loss approximation;
|
|
- AOF rewrite;
|
|
- disk full/write error where CI environment supports;
|
|
- declared data-loss/RPO evidence;
|
|
- restore reconciliation;
|
|
- session security epoch;
|
|
- fencing high watermark.
|
|
|
|
### 37.18 network fault matrix
|
|
|
|
Toxiproxy/netem/process control로:
|
|
|
|
| Fault | Expected |
|
|
| --- | --- |
|
|
| connect refused | known unavailable before send |
|
|
| latency > command timeout | read timeout or mutation indeterminate |
|
|
| response-only cut | applied-but-response-lost path |
|
|
| reconnect queue full | bounded reject, no heap growth |
|
|
| half-open connection | deadline/health transition |
|
|
| primary partition | failover semantics, no strong claim |
|
|
| DNS/endpoint change | re-resolution/rediscovery |
|
|
| TLS rotation | controlled reconnect |
|
|
| program busy | bounded failure/readiness |
|
|
|
|
### 37.19 program failure
|
|
|
|
- `SCRIPT FLUSH`;
|
|
- `NOSCRIPT`;
|
|
- wrong Function digest;
|
|
- missing library on one Cluster primary;
|
|
- result schema mismatch;
|
|
- malformed/wrong-type state;
|
|
- maximum argument;
|
|
- slow bounded program;
|
|
- intentionally long script in isolated test -> BUSY/alert/recovery;
|
|
- deployment rollback.
|
|
|
|
무한 script를 shared CI Redis에서 실행해 worker를 영구 block하지 않는다. isolated disposable
|
|
container와 hard timeout을 사용한다.
|
|
|
|
### 37.20 compatibility matrix
|
|
|
|
minimum:
|
|
|
|
- selected minimum Redis version;
|
|
- next supported minor;
|
|
- current approved major;
|
|
- managed-service compatible mode;
|
|
- Function mode;
|
|
- EVALSHA mode.
|
|
|
|
matrix:
|
|
|
|
- Spring Data Redis/Lettuce;
|
|
- program command set;
|
|
- key/codec result schema;
|
|
- session serializer/security version;
|
|
- topology.
|
|
|
|
“latest” floating image를 release gate에 사용하지 않는다. digest/version을 pin한다.
|
|
|
|
### 37.21 performance/capacity
|
|
|
|
benchmark pass/fail을 가짜 universal TPS로 고정하지 않는다. regression suite는 동일 controlled
|
|
environment에서:
|
|
|
|
- p50/p95/p99 operation duration;
|
|
- event-loop utilization;
|
|
- queue/in-flight;
|
|
- program server execution;
|
|
- memory per key;
|
|
- source fallback;
|
|
- hot-key throughput;
|
|
- failover recovery
|
|
|
|
를 baseline 대비 비교한다.
|
|
|
|
### 37.22 CI task
|
|
|
|
구현 시 아래 task명을 그대로 Gradle 공개 계약으로 만든다.
|
|
|
|
```text
|
|
:application-core:redisPolicyContractTest
|
|
:shared-contract:edgeRateLimitContractTest
|
|
:adapter:outbound:cache-redis:test
|
|
:adapter:outbound:cache-redis:redisStandaloneTest
|
|
:adapter:outbound:cache-redis:redisSecurityTest
|
|
:adapter:outbound:cache-redis:redisSentinelTest
|
|
:adapter:outbound:cache-redis:redisClusterTest
|
|
:adapter:outbound:cache-redis:redisFaultTest
|
|
:adapter:outbound:cache-redis:redisCompatibilityTest
|
|
:adapter:outbound:cache-redis:redisCacheCapabilityTest
|
|
:adapter:outbound:cache-redis:redisRateLimitCapabilityTest
|
|
:adapter:outbound:cache-redis:redisIdempotencyCapabilityTest
|
|
:adapter:outbound:cache-redis:redisSoftLeaseCapabilityTest
|
|
:adapter:outbound:cache-redis:redisFencedCoordinationCapabilityTest
|
|
:adapter:outbound:cache-redis:redisSessionCapabilityTest
|
|
:app-bootstrap:redisCompositionTest
|
|
redisCacheReadiness
|
|
redisRateLimitReadiness
|
|
redisIdempotencyReadiness
|
|
redisSoftLeaseReadiness
|
|
redisFencedCoordinationReadiness
|
|
redisSessionReadiness
|
|
redisProductionReadiness
|
|
redisAllImplementedCandidates
|
|
```
|
|
|
|
Redis leaf에는 `redisTest` source set을 만들고 source는
|
|
`src/redisTest/java`, resource는 `src/redisTest/resources`에 둔다. 여섯 `redis*Test` task는
|
|
동일 compiled source set에서 topology/evidence JUnit tag
|
|
`redis-standalone|redis-security|redis-sentinel|redis-cluster|redis-fault|redis-compatibility`와
|
|
card tag
|
|
`card-redis-cache|card-redis-edge-rate-limit|card-redis-request-replay-idempotency|
|
|
card-redis-cache-refresh-soft-lease|
|
|
card-redis-fenced-coordination|card-redis-session`을 함께 사용한다.
|
|
선택된 card/evidence tag expression 결과가 0개면 실패하며 Docker/service 부재도 readiness lane에서
|
|
skip하지 않는다.
|
|
application/shared/bootstrap의 contract task는 각 module의 별도
|
|
`src/redisPolicyContractTest`, `src/edgeRateLimitContractTest`, `src/redisCompositionTest` source
|
|
set을 사용해 일반 unit test와 production-readiness evidence를 구분한다.
|
|
|
|
card/evidence SSOT는 `src/config/redis/readiness-cards.yaml`로 고정한다.
|
|
|
|
canonical card ID와 Gradle task mapping:
|
|
|
|
| Card ID | Readiness task |
|
|
| --- | --- |
|
|
| `redis-cache` | `redisCacheReadiness` |
|
|
| `redis-edge-rate-limit` | `redisRateLimitReadiness` |
|
|
| `redis-request-replay-idempotency` | `redisIdempotencyReadiness` |
|
|
| `redis-cache-refresh-soft-lease` | `redisSoftLeaseReadiness` |
|
|
| `redis-fenced-coordination` | `redisFencedCoordinationReadiness` |
|
|
| `redis-session` | `redisSessionReadiness` |
|
|
|
|
registry key, capability descriptor ID, `card-<id>` tag, evidence artifact의 card ID는 이 표와 byte-for-byte
|
|
같아야 한다. short alias를 허용하지 않는다.
|
|
|
|
현재 card 상태와 topology/evidence는 이 문서에 복제하지 않으며
|
|
`src/config/redis/readiness-cards.yaml`만을 따른다. 현재 `selected` card는 없으며,
|
|
`implemented-candidate`는 release selection 또는 R2 qualification을 뜻하지 않는다.
|
|
|
|
각 `redis<Card>Readiness` task는 이 registry의 해당 card tag와 required evidence tag의 교집합을
|
|
실행하고, category마다 test count > 0, 성공 artifact, image/program/config digest를 요구한다.
|
|
`selected-topology`는 registry의 exact topology tag로 치환한다. 다른 card의 test가 대신 evidence를
|
|
채울 수 없다.
|
|
|
|
task dependency는 다음으로 고정한다.
|
|
|
|
- `:adapter:outbound:cache-redis:check` ->
|
|
`:adapter:outbound:cache-redis:redisStandaloneTest`;
|
|
- 각 card readiness -> 해당 capability test + required evidence/topology filtered test;
|
|
- root `redisProductionReadiness` -> registry에서 `state=selected`인 card readiness만;
|
|
- root `redisProductionReadiness` ->
|
|
application/shared contract, bootstrap composition,
|
|
`verifyCleanArchitectureDependencies`, `verifyEnvKeys`, `verifyPublicPathSnapshot`,
|
|
`verifyConfigurationPropertiesProcessor`;
|
|
- selected card가 0개면 provider-disabled/zero-side-effect composition과 no-false-R2 descriptor를
|
|
검증하고 모든 card를 `not selected`로 보고하며 real Redis task를 가장해 실행하지 않음;
|
|
- `redisAllImplementedCandidates` -> `selected`와 `implemented-candidate` card 전체를 nightly
|
|
실행하되 release label을 바꾸지 않음;
|
|
- release workflow는 root `redisProductionReadiness` 하나만 호출해 gate 누락을 막는다.
|
|
|
|
CLI `-PredisCards=`로 release 선택을 바꿀 수 없고 checked-in registry와 release profile digest만
|
|
selection authority다. `not-implemented` card의 test/tag가 0개인 것은 failure가 아니라
|
|
`not selected`; selected card의 missing test/evidence만 failure다.
|
|
|
|
server/container image SSOT는 `src/gradle/redis-test-images.properties`다. 최소 다음 key를
|
|
version control한다.
|
|
|
|
```properties
|
|
redis.minimum.image=<registry>/<image>:<exact-version>@sha256:<digest>
|
|
redis.next-minor.image=<registry>/<image>:<exact-version>@sha256:<digest>
|
|
redis.approved.image=<registry>/<image>:<exact-version>@sha256:<digest>
|
|
toxiproxy.image=<registry>/<image>:<exact-version>@sha256:<digest>
|
|
```
|
|
|
|
tag만 있거나 digest가 없거나 placeholder/`latest`이면 configuration 단계에서 실패한다.
|
|
Sentinel/Cluster container도 이 Redis image를 재사용하며 test resource config의 server version과
|
|
manifest minimum version이 불일치하면 compatibility task가 실패한다.
|
|
|
|
### 37.23 CI lane
|
|
|
|
`.github/workflows/ci-quality-gates.yml`의 PR blocking job `redis-standalone`:
|
|
|
|
- `:application-core:redisPolicyContractTest`;
|
|
- `:shared-contract:edgeRateLimitContractTest`;
|
|
- `:adapter:outbound:cache-redis:check`;
|
|
- `:app-bootstrap:redisCompositionTest`;
|
|
- 네 기존 gate
|
|
`verifyCleanArchitectureDependencies`, `verifyEnvKeys`, `verifyPublicPathSnapshot`,
|
|
`verifyConfigurationPropertiesProcessor`.
|
|
|
|
현재 branch protection이 단일 `release-gate` 집계 job을 required check로 사용하므로 workflow의
|
|
집계 계약도 함께 바꾼다.
|
|
|
|
```yaml
|
|
release-gate:
|
|
needs:
|
|
- quality-gates
|
|
- sample-off
|
|
- gate-matrix-lint
|
|
- redis-standalone
|
|
```
|
|
|
|
`Require every current blocking job to succeed` step에
|
|
`REDIS_RESULT: ${{ needs.redis-standalone.result }}`를 추가하고 기존 result loop가 이 값도
|
|
`success`인지 검사한다. job만 추가하고 `release-gate.needs`/검사 loop를 바꾸지 않은 상태는
|
|
PR blocking으로 인정하지 않는다. workflow contract test는 blocking job set과 aggregator
|
|
`needs`/env/result-check set이 정확히 같은지 검증한다.
|
|
|
|
새 `.github/workflows/redis-production-readiness.yml`은 `schedule`, `workflow_dispatch`,
|
|
release candidate trigger를 받는다. 첫 `resolve-redis-readiness` job이 checked-in card registry를
|
|
검증하고 selected/implemented-candidate card와 task/digest matrix를 artifact/output으로 만든다.
|
|
|
|
다음 topology job은 nightly의 공통 runtime/implemented-candidate qualification이다.
|
|
|
|
| Job | Required task |
|
|
| --- | --- |
|
|
| `redis-security` | `redisSecurityTest` |
|
|
| `redis-sentinel` | `redisSentinelTest` |
|
|
| `redis-cluster` | `redisClusterTest` |
|
|
| `redis-fault` | `redisFaultTest` |
|
|
| `redis-compatibility` | `redisCompatibilityTest` |
|
|
|
|
release candidate에서는 `selected-card-readiness` matrix가 selected card별
|
|
`redis<Card>Readiness`를 병렬 실행한다. card가 요구하지 않는 Sentinel/Cluster나 미구현 capability
|
|
task는 release dependency가 아니다. 최종 `redis-production-readiness` job은 selected card
|
|
matrix가 모두 성공한 뒤 fresh runner에서 root `redisProductionReadiness`를 다시 실행해 registry
|
|
digest와 evidence artifact set을 대조한다. selected production topology, TLS/ACL, Function digest,
|
|
image/license, recovery/runbook drill evidence가 없으면 해당 card의 release readiness가 아니다.
|
|
|
|
nightly `redis-all-candidates`는 `redisAllImplementedCandidates`를 실행한다. 그 실패는 candidate
|
|
품질 신호/승격 blocker지만 현재 selected card의 이미 존재하는 release evidence를 다른 card
|
|
미구현 때문에 자동 취소하지 않는다.
|
|
|
|
각 job은 `build/redis-evidence` 아래에서 allowlist schema로 다시 생성한 bounded manifest,
|
|
capability card, sanitized test summary만 artifact로 올린다. Gradle의 raw JUnit XML/HTML,
|
|
`system-out`/`system-err`, stack trace, container log/inspect, TLS/ACL fixture material은 업로드하지
|
|
않는다. 실제 사용 image attestation과 실제 topology/fault event chain을 수집하지 못한 현재
|
|
artifact는 각각 `NOT_CAPTURED`와 `releaseQualification=NOT_CLAIMED`를 기록하며, reconciler는
|
|
이 상태의 future `selected` 승격을 실패시킨다. secret/reference value, raw endpoint/key/value,
|
|
session/idempotency/lease token은 artifact에 포함하지 않는다. Candidate artifact retention은
|
|
짧게, 실제 release evidence는 조직의 audit retention 정책에 맞춘다.
|
|
|
|
### 37.24 no silent skip
|
|
|
|
developer local `test`는 Docker absence에서 explicit skipped report를 허용할 수 있다. 그러나
|
|
`redisProductionReadiness`와 release CI는 selected card의 required evidence service unavailable을
|
|
failure로 처리한다. unselected card는 skip이 아니라 `not selected`다.
|
|
|
|
report에는:
|
|
|
|
```text
|
|
executed
|
|
skipped with reason
|
|
not selected
|
|
failed
|
|
```
|
|
|
|
를 구분한다.
|
|
|
|
## 38. Gradle, dependency, artifact design
|
|
|
|
### 38.1 production dependency
|
|
|
|
Redis leaf target:
|
|
|
|
```groovy
|
|
dependencies {
|
|
implementation project(':application-core')
|
|
implementation project(':shared-contract')
|
|
implementation project(':adapter:outbound:support')
|
|
|
|
implementation 'org.springframework.boot:spring-boot-autoconfigure'
|
|
implementation 'org.springframework.data:spring-data-redis'
|
|
implementation 'io.lettuce:lettuce-core'
|
|
implementation 'org.slf4j:slf4j-api'
|
|
|
|
// 실제 session repository provider를 이 leaf가 소유할 때만 둘 다 직접 선언
|
|
implementation 'org.springframework.session:spring-session-core'
|
|
implementation 'org.springframework.session:spring-session-data-redis'
|
|
|
|
annotationProcessor 'org.springframework.boot:spring-boot-configuration-processor'
|
|
}
|
|
```
|
|
|
|
위 artifact와 ownership을 구현 계약으로 고정하고 호환 version은 Spring Boot 4 BOM과 lockfile이
|
|
결정한다. dependency report는 runtime client가 정확히 하나인지와 direct ownership을 검증한다.
|
|
`:adapter:inbound:web`과 `:app-bootstrap`의 조건부 `spring-session-core` 직접 소유, Redis leaf의
|
|
Servlet 금지, explicit auto-configuration suppression은 §31.3 표와 contract test를 따른다.
|
|
|
|
### 38.2 starter policy
|
|
|
|
application-core에 starter를 추가하지 않는다. Redis leaf도 broad starter가 불필요한
|
|
auto-configuration/connection을 만들지 않도록 direct dependency와 explicit configuration을
|
|
선호한다.
|
|
|
|
starter를 사용하더라도:
|
|
|
|
- capability disabled 시 side effect 없음;
|
|
- one global factory auto-created 안 됨;
|
|
- role별 qualifier;
|
|
- app-bootstrap이 composition owner
|
|
|
|
를 test해야 한다.
|
|
|
|
### 38.3 test dependency
|
|
|
|
```text
|
|
Testcontainers JUnit
|
|
Testcontainers core/GenericContainer
|
|
Toxiproxy module
|
|
AssertJ/JUnit
|
|
property-test library already approved by repository
|
|
```
|
|
|
|
Redis-specific unofficial embedded server를 semantic evidence로 사용하지 않는다.
|
|
|
|
Sentinel/Cluster config asset과 container orchestration은 test resources/scripts에 versioned로 둔다.
|
|
|
|
### 38.4 dependency boundary
|
|
|
|
architecture gate:
|
|
|
|
- Redis/Spring Data/Lettuce/Session import는 Redis leaf/bootstrap/web의 approved package만;
|
|
- `domain-core`, `application-core`, `shared-contract`에는 없음;
|
|
- inbound web은 outbound Redis package에 의존하지 않음;
|
|
- JPA와 Redis provider leaf 간 직접 dependency 없음;
|
|
- sample production 역의존 없음.
|
|
|
|
### 38.5 version baseline
|
|
|
|
R2 portable minimum은 Redis 7.2로 고정한다. 실제 product image는 이 minimum 이상인 별도
|
|
승인 exact version/digest로 고정한다.
|
|
|
|
이유:
|
|
|
|
- Redis Functions가 존재하는 generation;
|
|
- `WAITAOF` capability negotiation 가능;
|
|
- newer managed versions에서도 legacy Lua primitive 사용 가능.
|
|
|
|
새er Redis command에 맞춰 baseline을 몰래 올리지 않는다. 예를 들어 newer conditional delete/set
|
|
command가 있어도 minimum matrix가 지원하기 전에는 owner-safe Function/Lua 구현을 유지한다.
|
|
|
|
정확한 approved server image와 EOL/support policy는 product ADR에서 결정한다.
|
|
|
|
### 38.6 license/image gate
|
|
|
|
Redis release line에 따라 license 선택지가 달라질 수 있다.
|
|
|
|
- image source;
|
|
- exact version/digest;
|
|
- organization legal approval;
|
|
- managed provider terms;
|
|
- vulnerability/EOL;
|
|
- upgrade/rollback
|
|
|
|
을 ADR/CI metadata에 기록한다. floating `redis:latest`를 production/readiness evidence로 쓰지
|
|
않는다.
|
|
|
|
### 38.7 dependency lock
|
|
|
|
새 production/test dependency 추가 시:
|
|
|
|
- Boot BOM ownership 확인;
|
|
- BOM 밖 dependency version SSOT;
|
|
- `resolveAndLockAll --write-locks`;
|
|
- strict lock verification;
|
|
- runtimeClasspath에서 client exactly one;
|
|
- duplicate Netty/client conflicts;
|
|
- license/SBOM/security scan.
|
|
|
|
### 38.8 program artifact
|
|
|
|
Function/Lua는 JAR resource와 별도 deployable manifest를 함께 만든다.
|
|
|
|
```text
|
|
program-set.json
|
|
functions/*.lua
|
|
scripts/*.lua
|
|
checksums.sha256
|
|
compatibility.json
|
|
```
|
|
|
|
JAR implementation version/source revision과 program manifest를 연결한다.
|
|
|
|
### 38.9 capability card artifact
|
|
|
|
build가 machine-readable card를 생성한다.
|
|
|
|
```text
|
|
provider IDs
|
|
readiness
|
|
roles
|
|
minimum Redis version
|
|
program set/digest
|
|
key/codec versions
|
|
guarantees/non-guarantees
|
|
required settings
|
|
test evidence profile
|
|
```
|
|
|
|
startup과 docs가 서로 다른 수동 목록을 유지하지 않게 한다.
|
|
|
|
## 39. Implementation migration
|
|
|
|
readiness는 Redis leaf 전체에 한 번에 부여하지 않고 capability card별로 승격한다.
|
|
|
|
| Capability card | R2 최소 구현/evidence | 예정 phase |
|
|
| --- | --- | --- |
|
|
| `redis-cache` | standalone real provider, key/codec/TTL, soft/hard/stale, generation/invalidation, source bulkhead, outage/memory/security test | 1 + 2 + 5 |
|
|
| `redis-edge-rate-limit` | typed outcome, 선택 알고리즘 golden/property/concurrency/fault test | 2 + 5 |
|
|
| `redis-request-replay-idempotency` | owner-safe state machine, replay/lease TTL, JPA cutover, response-loss reconciliation | 3 + 5 |
|
|
| `redis-cache-refresh-soft-lease` | cache refresh integration, owner-safe renew/release/lost, duplicate-holder 허용 계약, fault test | 2 + 3 + 5 |
|
|
| `redis-fenced-coordination` | monotonic token과 protected-resource stale-token rejection evidence | 3 + 5, 선택 시에만 |
|
|
| `redis-session` | versioned repository, cookie/CSRF/fixation, multi-pod/logout race/failover test | 4 + 5 |
|
|
|
|
공통 runtime가 R2 evidence를 가져도 미구현 capability는 R0이고, cache가 R2여도 session/rate가
|
|
자동으로 R2가 되지 않는다. Phase 5는 각 card가 요구하는 selected topology/security/fault
|
|
evidence를 따로 묶어 승격한다. §37.22 registry의 해당 card가 `selected`이고 exact
|
|
`redis<Card>Readiness` evidence digest가 성공한 경우에만 descriptor를 R2로 바꾼다.
|
|
|
|
### Phase 0 — current truth와 contract freeze
|
|
|
|
- 현재 R0 seam/limitation을 README에 정직하게 표시;
|
|
- `APP_CACHE_REDIS_ENABLED`가 real client 없이 실패함을 문서화;
|
|
- current focused test 유지;
|
|
- cache/rate/idempotency/lease/session semantic contract 승인;
|
|
- capability descriptor/readiness/failure outcome;
|
|
- module/package migration ADR;
|
|
- implementation plan 작성.
|
|
|
|
Acceptance:
|
|
|
|
- 범용 Redis port 없음;
|
|
- current production-ready 오표기 없음;
|
|
- core dependency 방향 승인.
|
|
|
|
### Phase 1 — runtime, key, codec, cache R1 / R2 foundation
|
|
|
|
- Spring Data Redis + Lettuce direct dependency;
|
|
- standalone typed runtime;
|
|
- role/deployment config;
|
|
- key builder/HMAC/hash slot;
|
|
- codec/envelope;
|
|
- program registry;
|
|
- cache region port와 cache-aside executor;
|
|
- TTL/negative/invalidate/jitter;
|
|
- local single-flight/source bulkhead;
|
|
- real standalone integration;
|
|
- health/metrics/security baseline.
|
|
|
|
Acceptance:
|
|
|
|
- 실제 Redis cache provider 동작;
|
|
- disabled zero side effect;
|
|
- miss/unavailable/corrupt 구분;
|
|
- bounded TTL/payload/source load;
|
|
- no SDK in core;
|
|
- focused/architecture/readiness test 통과.
|
|
|
|
### Phase 2 — advanced cache와 distributed rate
|
|
|
|
- soft/hard TTL/stale;
|
|
- generation/revision invalidation;
|
|
- distributed refresh lease;
|
|
- edge rate shared contract;
|
|
- fixed/sliding counter/token bucket;
|
|
- policy registry/revision;
|
|
- evaluation dedup;
|
|
- local emergency failure policy;
|
|
- inbound HTTP mapping migration;
|
|
- current unbounded local map 제거/안전 fallback화.
|
|
|
|
Acceptance:
|
|
|
|
- multi-pod quota contract;
|
|
- algorithm concurrency/property;
|
|
- correct `Retry-After`;
|
|
- raw principal/IP key 없음;
|
|
- bean-name multi-instance validation 제거.
|
|
|
|
### Phase 3 — owner-safe idempotency와 lease
|
|
|
|
- idempotency port v2;
|
|
- JPA provider migration;
|
|
- Redis state machine;
|
|
- processing/replay TTL 분리;
|
|
- owner-safe complete/release;
|
|
- indeterminate reconciliation;
|
|
- lease v2/renew/lost;
|
|
- Redis efficiency provider;
|
|
- provider selector/cutover runbook;
|
|
- optional fenced contract와 resource fixture.
|
|
|
|
Acceptance:
|
|
|
|
- stale owner mutation 차단;
|
|
- cross-store non-guarantee 명시;
|
|
- failover/fault contract;
|
|
- JDBC/Redis provider 혼합 activation 없음.
|
|
|
|
### Phase 4 — Redis Session
|
|
|
|
- exclusive `jwt|redis-session`;
|
|
- session role/deployment;
|
|
- Spring Session repository;
|
|
- explicit serializer;
|
|
- cookie/CSRF/fixation;
|
|
- idle/absolute expiry;
|
|
- concurrent logout/save;
|
|
- multi-pod test;
|
|
- fail closed/readiness;
|
|
- indexed repository는 별도 opt-in.
|
|
|
|
Acceptance:
|
|
|
|
- JWT mode Redis side effect 0;
|
|
- session mode multi-pod/security/serializer/failure test 통과;
|
|
- cache deployment와 물리 격리.
|
|
|
|
### Phase 5 — capability별 topology/security/failure R2 promotion
|
|
|
|
- Sentinel profile;
|
|
- Cluster-compatible program/key;
|
|
- TLS/ACL;
|
|
- credential rotation;
|
|
- queue/backpressure/reconnect certainty;
|
|
- memory/eviction/persistence;
|
|
- fault matrix;
|
|
- production-readiness CI;
|
|
- runbook.
|
|
|
|
Acceptance:
|
|
|
|
- 승격 대상 capability마다 selected production topology R2 evidence;
|
|
- capability card별 evidence bundle과 readiness label;
|
|
- 증거가 없는 capability는 R0/R1 유지;
|
|
- no silent skip;
|
|
- program/ACL/schema conformance.
|
|
|
|
#### Phase 5A — Sentinel-first R2 qualification slice
|
|
|
|
Phase 5 전체를 한 번에 구현하지 않는다. 먼저 §37.13의 disposable 3-node k3s Sentinel 환경에서
|
|
다음 순서로 진행한다.
|
|
|
|
1. lab lifecycle/preflight/host-isolation contract를 테스트 우선으로 고정한다;
|
|
2. Sentinel discovery와 Redis data runtime을 별도 auth/trust/lifecycle로 구현한다;
|
|
3. quorum-consistent discovery, bounded rediscovery, qualified runtime swap와 bounded drain을
|
|
구현한다;
|
|
4. security positive/negative test 후 rate limit, idempotency, session의 multi-pod 정상 경로를
|
|
실행한다;
|
|
5. primary kill과 response-loss fault를 주입하고 capability invariant와 `INDETERMINATE`
|
|
semantics를 검증한다;
|
|
6. image/fault timeline을 실제 관측에서 생성하고 sanitizer/reconciler를 통과시킨다;
|
|
7. focused/full Gradle verification과 독립 review를 마친 뒤 이 slice에서 멈춘다.
|
|
|
|
이번 slice에 포함하지 않는 항목:
|
|
|
|
- Redis Cluster와 Cluster cache scaling;
|
|
- fenced coordination;
|
|
- R3 capacity soak/long chaos/reshard;
|
|
- k3s control-plane HA, physical host/AZ failure;
|
|
- full credential/certificate rotation drill;
|
|
- optional cache의 Sentinel release promotion.
|
|
|
|
이번 slice의 agent-side 종료 상태는 `R2-ready candidate`다. repository가 human-only commit
|
|
policy를 사용하므로 clean committed source와 실제 remote GitHub Actions evidence는 사람이
|
|
수행하는 최종 promotion gate다. 이 두 증거가 없으면 readiness card를 `selected`로 바꾸거나
|
|
R2라고 표시하지 않는다.
|
|
|
|
### Phase 6 — R3와 split review
|
|
|
|
- actual Cluster reshard/failover;
|
|
- rolling serializer/program/key upgrade;
|
|
- capacity soak;
|
|
- restore drill;
|
|
- fencing high-watermark recovery;
|
|
- multi-region 필요성 검토;
|
|
- capability별 leaf split trigger 재평가;
|
|
- external platform artifact 추출 검토.
|
|
|
|
## 40. 완료 기준
|
|
|
|
특정 Redis capability card가 R2라고 주장하려면 아래 공통 조건과 §39의 해당 card 조건을 모두
|
|
충족해야 한다. “Redis module 전체 R2”라는 단일 label은 사용하지 않는다.
|
|
|
|
- real Redis client/provider가 있음;
|
|
- enabled capability가 실제 consumer port와 연결됨;
|
|
- disabled capability side effect 0;
|
|
- role별 deployment 분리;
|
|
- typed topology/TLS/ACL/timeout/queue setting;
|
|
- cache/idempotency/lease/session/rate failure policy 분리;
|
|
- key namespace/HMAC/version/hash-slot;
|
|
- explicit codec, no JDK serialization;
|
|
- bounded payload/collection/cardinality;
|
|
- value+TTL atomic write;
|
|
- versioned bounded program catalog;
|
|
- Function/Lua deployment/digest/recovery;
|
|
- mutation `INDETERMINATE` outcome;
|
|
- cache miss/unavailable/corrupt 분리;
|
|
- cache-aside/source bulkhead/stampede defense;
|
|
- policy별 rate algorithm과 fallback;
|
|
- owner-safe lease renew/release;
|
|
- idempotency owner token과 processing/replay TTL 분리;
|
|
- session cookie/CSRF/fixation/serializer/multi-pod;
|
|
- no cross-store exactly-once claim;
|
|
- no strong Redis lock claim;
|
|
- memory/eviction/persistence attestation;
|
|
- liveness/readiness 분리;
|
|
- bounded metrics/log/trace;
|
|
- graceful lifecycle;
|
|
- standalone real-service test;
|
|
- selected topology/failure/security test;
|
|
- dependency/architecture/env/public path gate;
|
|
- runbook/capability card;
|
|
- LLM Wiki capture.
|
|
|
|
### 40.2 Sentinel-first slice 종료 게이트
|
|
|
|
§37.13과 Phase 5A의 작업은 아래가 모두 충족된 경우에만 `R2-ready candidate`로 종료한다.
|
|
|
|
- exact VM inventory와 dedicated kubeconfig로 lab create/verify/destroy가 반복 가능하다;
|
|
- host k3s context, node, workload와 default kubeconfig의 전/후 fingerprint가 같다;
|
|
- Sentinel discovery와 Redis data auth/trust가 분리되고 negative security test가 통과한다;
|
|
- primary kill 뒤 quorum election, qualified runtime swap, bounded drain과 semantic readiness
|
|
recovery의 실제 timeline이 있다;
|
|
- rate limit, idempotency, session을 서로 다른 pod에서 검증하고 failover 뒤 invariant가
|
|
유지된다;
|
|
- confirmed acknowledgement와 `INDETERMINATE`를 구분하며 blind mutation replay가 없다;
|
|
- actual image/config/program digest와 sanitized evidence가 reconciler를 통과한다;
|
|
- focused test, Redis readiness 관련 task, repository `test`/`check`, architecture/env/public-path
|
|
gate와 독립 review가 통과한다;
|
|
- exact allowlist VM teardown과 lab resource 정리 결과가 기록된다.
|
|
|
|
위 조건은 clean committed source와 실제 remote CI를 대신하지 않는다. 두 최종 promotion
|
|
증거가 없으면 card 상태는 `implemented-candidate`, `releaseQualification=NOT_CLAIMED`를
|
|
유지한다. 종료 뒤 Redis Cluster/R3/fenced coordination 또는 fileserver/HTTP client로 자동으로
|
|
넘어가지 않고 다음 우선순위를 다시 결정한다.
|
|
|
|
R3는 추가로:
|
|
|
|
- failover/partition;
|
|
- Cluster reshard;
|
|
- rolling compatibility;
|
|
- recovery drill;
|
|
- capacity/latency evidence;
|
|
- credential/certificate rotation;
|
|
- program deployment rollback
|
|
|
|
을 실제 topology에서 증명해야 한다.
|
|
|
|
### 40.1 금지 문구
|
|
|
|
- “Redis는 single-thread라 multi-command도 race가 없다.”
|
|
- “Lua를 쓰므로 Redis 전체 성능에 영향이 없다.”
|
|
- “Lua/Function이므로 cross-store transaction이다.”
|
|
- “AOF와 replica가 있으므로 write loss가 없다.”
|
|
- “`WAIT`를 호출하므로 strong consistency다.”
|
|
- “Redis lock을 썼으므로 correctness가 보장된다.”
|
|
- “Redlock이면 fencing이 필요 없다.”
|
|
- “idempotency key가 있으므로 side effect가 exactly once다.”
|
|
- “Pub/Sub invalidation이 있으므로 stale cache가 없다.”
|
|
- “keyspace notification이 exact expiry event다.”
|
|
- “database number를 나눴으므로 session/cache가 격리됐다.”
|
|
- “PING이 성공하므로 Redis capability가 healthy다.”
|
|
- “timeout이므로 command는 실행되지 않았다.”
|
|
- “pipeline이므로 atomic하다.”
|
|
- “Spring Session을 추가했으므로 secure session이다.”
|
|
- “Testcontainers standalone이 통과했으므로 Cluster/HA도 production-ready다.”
|
|
|
|
## 41. 운영 runbook 요구
|
|
|
|
각 항목은 detection, immediate mitigation, safety decision, recovery, verification을 포함한다.
|
|
|
|
### Runtime/topology
|
|
|
|
- connection/auth/TLS failure;
|
|
- DNS/managed endpoint change;
|
|
- Sentinel failover;
|
|
- Cluster `MOVED`/`ASK` storm;
|
|
- uncovered slot;
|
|
- advertised node unreachable;
|
|
- reconnect queue saturation;
|
|
- event-loop/connection exhaustion;
|
|
- rolling client upgrade.
|
|
|
|
### Program/schema
|
|
|
|
- Function missing/digest mismatch;
|
|
- `NOSCRIPT`;
|
|
- BUSY/slow script;
|
|
- result schema mismatch;
|
|
- key schema rolling migration;
|
|
- codec corrupt/future version;
|
|
- program rollback;
|
|
- missing program on one Cluster primary.
|
|
|
|
### Cache
|
|
|
|
- hit-ratio collapse;
|
|
- source load storm;
|
|
- hot/big key;
|
|
- mass invalidation;
|
|
- stale data incident;
|
|
- generation key loss;
|
|
- negative cache abuse;
|
|
- cache warmup/cold restart;
|
|
- L1 invalidation disconnect.
|
|
|
|
### Rate limit
|
|
|
|
- Redis unavailable;
|
|
- fail-closed incident;
|
|
- local emergency activation;
|
|
- hot global policy key;
|
|
- incorrect policy revision;
|
|
- evaluation double charge;
|
|
- algorithm migration/shadow;
|
|
- subject cardinality attack.
|
|
|
|
### Lease/fencing
|
|
|
|
- lease renewal loss;
|
|
- duplicate holder after failover;
|
|
- stale owner release;
|
|
- fencing token regression;
|
|
- protected resource rejection spike;
|
|
- leader task cancellation;
|
|
- semaphore permit leak;
|
|
- high-watermark repair.
|
|
|
|
### Idempotency
|
|
|
|
- stuck in-progress;
|
|
- owner takeover;
|
|
- complete indeterminate;
|
|
- fingerprint mismatch spike;
|
|
- response corruption/oversize;
|
|
- Redis record loss after business commit;
|
|
- JDBC/Redis provider cutover;
|
|
- manual reconciliation/abandonment.
|
|
|
|
### Session
|
|
|
|
- repository outage/re-auth spike;
|
|
- serializer incompatibility;
|
|
- session resurrection;
|
|
- mass logout/revoke;
|
|
- key/index orphan;
|
|
- absolute/idle expiry drift;
|
|
- backup restore security epoch;
|
|
- credential rotation;
|
|
- Cluster indexed repository cleanup.
|
|
|
|
### Memory/persistence
|
|
|
|
- cache eviction spike;
|
|
- coordination/session eviction;
|
|
- noeviction OOM;
|
|
- fragmentation/RSS;
|
|
- replication backlog;
|
|
- AOF/RDB failure;
|
|
- disk full;
|
|
- rewrite/fork latency;
|
|
- backup restore;
|
|
- capacity scale-out.
|
|
|
|
### Security
|
|
|
|
- credential compromise;
|
|
- ACL drift;
|
|
- unauthorized command attempt;
|
|
- certificate expiry/rotation;
|
|
- secret leakage in logs;
|
|
- unexpected public exposure;
|
|
- destructive operator command;
|
|
- Redis image/license/security update.
|
|
|
|
## 42. Primary references
|
|
|
|
### Redis execution and programmability
|
|
|
|
- [Redis Lua scripting, atomic blocking execution, key declaration, and script cache](https://redis.io/docs/latest/develop/programmability/eval-intro/)
|
|
- [Redis Functions](https://redis.io/docs/latest/develop/programmability/functions-intro/)
|
|
- [Redis transactions and `WATCH`](https://redis.io/docs/latest/develop/using-commands/transactions/)
|
|
- [Redis multi-key operations](https://redis.io/docs/latest/develop/using-commands/multi-key-operations/)
|
|
- [Redis latency and slow-command guidance](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/)
|
|
|
|
### Redis topology, durability, and memory
|
|
|
|
- [Redis Cluster specification](https://redis.io/docs/latest/operate/oss_and_stack/reference/cluster-spec/)
|
|
- [Redis Sentinel](https://redis.io/docs/latest/operate/oss_and_stack/management/sentinel/)
|
|
- [Redis replication](https://redis.io/docs/latest/operate/oss_and_stack/management/replication/)
|
|
- [`WAIT`](https://redis.io/docs/latest/commands/wait/)
|
|
- [`WAITAOF`](https://redis.io/docs/latest/commands/waitaof/)
|
|
- [Redis persistence](https://redis.io/docs/latest/operate/oss_and_stack/management/persistence/)
|
|
- [Redis key eviction](https://redis.io/docs/latest/develop/reference/eviction/)
|
|
- [`MEMORY USAGE`](https://redis.io/docs/latest/commands/memory-usage/)
|
|
- [`EXPIRE`](https://redis.io/docs/latest/commands/expire/)
|
|
- [Redis keyspace guidance and production `KEYS` warning](https://redis.io/docs/latest/develop/use/keyspace/)
|
|
|
|
### Redis capability patterns
|
|
|
|
- [Redis rate-limiter use case](https://redis.io/docs/latest/develop/use-cases/rate-limiter/)
|
|
- [Redis rate-limiter algorithm comparison](https://redis.io/tutorials/howtos/ratelimiting/)
|
|
- [Redis distributed lock pattern and limitations](https://redis.io/docs/latest/develop/clients/patterns/distributed-locks/)
|
|
- [Redis cache-aside](https://redis.io/docs/latest/develop/use-cases/cache-aside/)
|
|
- [Redis client-side caching](https://redis.io/docs/latest/develop/clients/client-side-caching/)
|
|
- [Redis session-store use case](https://redis.io/docs/latest/develop/use-cases/session-store/)
|
|
- [Redis Pub/Sub](https://redis.io/docs/latest/develop/pubsub/)
|
|
- [Redis keyspace notifications](https://redis.io/docs/latest/develop/pubsub/keyspace-notifications/)
|
|
- [Redis Streams](https://redis.io/docs/latest/develop/data-types/streams/)
|
|
|
|
### Security and operations
|
|
|
|
- [Redis security](https://redis.io/docs/latest/operate/oss_and_stack/management/security/)
|
|
- [Redis ACL](https://redis.io/docs/latest/operate/oss_and_stack/management/security/acl/)
|
|
- [Redis TLS](https://redis.io/docs/latest/operate/oss_and_stack/management/security/encryption/)
|
|
- [Redis latency monitoring](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency-monitor/)
|
|
- [Redis `SLOWLOG`](https://redis.io/docs/latest/commands/slowlog/)
|
|
- [Redis CLI key and hot-key inspection](https://redis.io/docs/latest/develop/tools/cli/)
|
|
- [Redis licenses](https://redis.io/legal/licenses/)
|
|
|
|
### Java client and Spring
|
|
|
|
- [Spring Data Redis reference](https://docs.spring.io/spring-data/redis/reference/)
|
|
- [Spring Data Redis drivers](https://docs.spring.io/spring-data/redis/reference/redis/drivers.html)
|
|
- [Spring Data Redis scripting](https://docs.spring.io/spring-data/redis/reference/redis/scripting.html)
|
|
- [Spring Data Redis transactions](https://docs.spring.io/spring-data/redis/reference/redis/transactions.html)
|
|
- [Spring Data Redis pipelining](https://docs.spring.io/spring-data/redis/reference/redis/pipelining.html)
|
|
- [Spring Data Redis serialization](https://docs.spring.io/spring-data/redis/reference/redis/template.html)
|
|
- [Spring Session repository APIs](https://docs.spring.io/spring-session/reference/api.html)
|
|
- [Spring Session Redis configuration and indexed-repository caveat](https://docs.spring.io/spring-session/reference/configuration/redis.html)
|
|
- [Lettuce command execution reliability](https://redis.github.io/lettuce/advanced-usage/command-execution-reliability/)
|
|
- [Lettuce client options](https://redis.github.io/lettuce/advanced-usage/client-options/)
|