# Redis Production Capability Completion Plan > **Scope:** Redis를 먼저 완료한다. 현재 실행 단위는 deep design Phase 5 전체가 아니라 > `Sentinel-first R2 qualification slice`다. 이 slice의 검증과 보고가 끝나면 멈추고 > fileserver, HTTP client, Redis Cluster/R3 중 다음 우선순위를 다시 정한다. > > **Workflow note:** 저장소가 지정한 Superpowers 설계·계획·TDD·디버깅·검증·리뷰 워크플로우를 > 적용한다. agent는 human-only commit 정책에 따라 stage/commit/amend/push하지 않는다. **Goal:** `2026-07-26-redis-production-capability-design.md`의 Phase 1–5를 capability별로 구현하고, standalone 기능의 존재를 production readiness로 오표기하지 않는 Redis platform을 만든다. **Architecture:** `application-core`와 `shared-contract`는 provider-neutral semantic contract만 소유한다. `adapter:outbound:cache-redis`가 Redis deployment, topology, key, codec, program, runtime과 capability provider를 소유한다. `adapter:inbound:web`은 HTTP rate/session 보안 매핑만, `app-bootstrap`은 provider/role/auth-mode composition만 소유한다. `domain-core`에는 Redis 개념을 추가하지 않는다. **Readiness rule:** Redis leaf 전체에 단일 R2 label을 부여하지 않는다. `redis-cache`, `redis-edge-rate-limit`, `redis-request-replay-idempotency`, `redis-cache-refresh-soft-lease`, `redis-fenced-coordination`, `redis-session` card가 독립적으로 승격한다. R3 증거가 없는 failover/reshard/rotation은 R2 범위로 과장하지 않는다. **Worktree rule:** 현재 `main` worktree의 다른 기술 변경은 사용자 소유다. Redis가 소유하지 않는 fileserver, HTTP client, messaging, notification, object storage 변경을 되돌리거나 포맷하지 않는다. **Current milestone exit:** agent-side 목표는 `R2-ready candidate`다. clean committed source와 실제 remote GitHub Actions evidence가 없으면 card를 `selected`로 바꾸거나 R2라고 주장하지 않는다. --- ## Task 0 — Baseline과 acceptance registry 고정 **Files** - Create: `src/config/redis/readiness-cards.yaml` - Create: `src/gradle/redis-test-images.properties` - Modify: `src/adapter/outbound/cache-redis/README.md` - Modify: `docs/superpowers/specs/2026-07-26-redis-production-capability-design.md` **Tests first** - registry가 canonical card ID 여섯 개를 정확히 한 번 포함하는지 실패 테스트를 작성한다. - image tag에 exact version과 digest가 없으면 configuration이 실패하는 테스트를 작성한다. - `selected`, `implemented-candidate`, `not-implemented` 이외 상태를 거절한다. - 현재 구현과 다른 readiness 표기를 거절한다. **Implementation** - 시작 상태는 cache/rate를 `implemented-candidate`, 나머지는 `not-implemented`로 기록한다. - 실제 required evidence가 생기기 전에는 어떤 card도 `selected` R2로 승격하지 않는다. - Redis minimum version은 실행 가능한 image/digest와 program manifest를 한 SSOT로 맞춘다. **Verification** ```bash cd src ./gradlew :adapter:outbound:cache-redis:test --tests '*RedisReadinessRegistryTest' --console=plain ``` ## Task 1 — Canonical deployment/topology/role model **Files** - Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderProperties.java` - Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettings.java` - Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactory.java` - Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRole.java` - Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRoleBinding.java` - Test: `src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactoryTest.java` - Test: `src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderPropertiesBindingTest.java` **Tests first** - topology는 `standalone|sentinel|cluster` 중 정확히 하나다. - endpoint는 non-empty, unique, bounded host/port다. - Sentinel은 master name, 최소 3개 discovery endpoint, data/Sentinel auth와 TLS를 분리한다. - Cluster는 database 0만 허용하고 seed가 비어 있으면 실패한다. - role은 존재하는 deployment만 참조한다. - cache와 session/coordination의 incompatible co-location을 startup 전에 거절한다. - provider 정의만 있고 capability binding이 없으면 runtime side effect가 0이다. **Implementation** - Spring binding class와 validated sealed runtime model을 분리한다. - legacy `app.cache.redis`와 `app.rate-limit`은 migration compiler 입력으로만 허용하고 canonical model과 동시에 설정되면 precedence를 정하지 않고 실패한다. - `ClientMode.EXTERNAL`을 topology로 취급하지 않는다. **Verification** ```bash cd src ./gradlew :adapter:outbound:cache-redis:test --tests '*RedisDeploymentSettings*' --console=plain ``` ## Task 2 — Topology-aware runtime, TLS/ACL과 secret material **Files** - Create: `.../redis/runtime/RedisDeploymentRuntime.java` - Create: `.../redis/runtime/RedisDeploymentRuntimeFactory.java` - Create: `.../redis/runtime/StandaloneRedisDeploymentRuntime.java` - Create: `.../redis/runtime/SentinelRedisDeploymentRuntime.java` - Create: `.../redis/runtime/ClusterRedisDeploymentRuntime.java` - Create: `.../redis/security/RedisCredentialMaterialProvider.java` - Create: `.../redis/security/RedisCredentialRotationCoordinator.java` - Modify: `src/adapter/outbound/cache-redis/build.gradle` - Modify: `src/adapter/outbound/cache-redis/gradle.lockfile` **Tests first** - standalone/Sentinel/Cluster가 각자 다른 native client/runtime을 만든다. - Sentinel discovery credential/trust와 data-node credential/trust가 섞이지 않는다. - Cluster client는 periodic+adaptive topology refresh, DB 0, bounded redirect/queue profile을 가진다. - production profile에서 plaintext, trust-all, hostname verification off를 거절한다. - named ACL username이 없거나 raw password가 YAML에 있으면 production activation이 실패한다. - duplicate/out-of-order rotation event, expiry 재조회, new connection 검증 실패가 old traffic을 안전하게 보존한다. - disabled capability는 client/event-loop/subscriber/scheduler를 만들지 않는다. **Implementation** - direct `spring-data-redis`, `lettuce-core` dependency를 leaf가 소유한다. - deployment별 client resources와 lifecycle을 소유한다. - connect/TLS/acquire/command/overall/shutdown timeout을 분리한다. - 기존 no-replay, disconnected reject, finite queue/count/byte admission을 topology runtime에도 보존한다. - secret value/reference/provider exception을 log/metric에 남기지 않는다. ## Task 3 — Key, codec, program manifest foundation **Files** - Create: `src/config/redis/program-set.schema.json` - Modify: `src/adapter/outbound/cache-redis/src/main/resources/redis/program-set.json` - Modify: `src/adapter/outbound/cache-redis/src/main/resources/redis/rate-program-set.json` - Modify: `.../redis/RedisProgramDescriptor.java` - Modify: `.../redis/RedisProgramCatalog.java` - Modify: `.../redis/RedisLuaProgramExecutor.java` - Create: `.../redis/key/RedisKeyMaterialProvider.java` - Create: `.../redis/codec/RedisCapabilityCodec.java` **Tests first** - 모든 program은 exact source digest, semantic version, ordered KEYS/ARGV, result schema, slot rule, state/TTL bound, minimum Redis version, retry/certainty, ACL command를 가진다. - manifest와 Java descriptor가 drift하면 build가 실패한다. - `NOSCRIPT` recovery는 bounded `SCRIPT LOAD -> EVALSHA`이고 arbitrary source 실행 surface가 없다. - same-resource multi-key는 real `CLUSTER KEYSLOT`과 같은 slot이다. - key digest material rotation은 fixed/dual-read-delete/cold-cutover rule을 지킨다. - cache/idempotency/session codec은 N/N-1, future/corrupt/oversize/forbidden type을 구분한다. **Implementation** - foundation/rate manifest를 하나의 versioned registry contract로 통합하되 capability package와 facade는 분리한다. - raw command, raw key, generic program executor를 Spring/application public surface에 노출하지 않는다. ## Task 4 — Cache consistency spine와 semantic region composition **Files** - Modify: `src/application-core/src/main/java/dev/caskeleton/application/cache/*` - Create: `.../redis/cache/RedisCacheGenerationStore.java` - Create: `.../redis/cache/RedisCacheRegionCompiler.java` - Add resources: `region-generation-init-v1.lua`, `region-generation-bump-v1.lua`, `cache-record-if-generation-v1.lua` - Modify: `.../redis/RedisStringCacheRegion.java` - Tests: application barrier tests, Redis real-service concurrency tests, binding tests **Tests first** - source load 중 generation bump가 일어나면 old result가 visible하지 않다. - captured generation과 source revision이 바뀌면 stale writer가 새 값을 덮어쓰지 않는다. - generation init race에서 하나의 canonical generation만 선택된다. - operation ID가 같은 bump replay는 한 번만 적용된다. - 여러 semantic region의 duplicate/missing binding은 fail-fast다. - 실제 consumer가 semantic `CacheRegionPort`와 `CacheAsideExecutor`를 사용하고 legacy fail-open router와 암묵적으로 섞이지 않는다. **Implementation decision** - source revision은 opaque하므로 lexical “newer” 비교를 하지 않는다. - region generation은 mass invalidation fence다. - per-key invalidation은 해당 key의 revision/tombstone fence를 사용해 region 전체를 bump하지 않는다. - write는 captured generation/revision condition을 만족할 때만 기록한다. ## Task 5 — Distributed refresh soft lease, L1/L2와 cache observability **Files** - Create application cache refresh coordination contracts without Redis types. - Create Redis refresh claim/release programs and semantic provider. - Create bounded L1 cache decorator and invalidation subscriber/reconciler. - Create framework-free cache observation events and Micrometer adapter instrumentation. - Update `docs/registries/metrics.yaml`. **Tests first** - 두 pod simulation에서 정상 시 refresh owner는 하나다. - lease expiry에서는 duplicate load를 허용하지만 generation guard가 stale write를 차단한다. - disconnected invalidation subscriber는 L1을 flush하고 generation을 재확인한다. - Pub/Sub event loss에도 L1 TTL/generation reconciliation으로 stale bound를 지킨다. - L1 max weight/cardinality/TTL, subscriber queue, refresh scheduler가 모두 bounded다. - Redis liveness는 애플리케이션 liveness를 내리지 않는다. - optional cache outage는 `DEGRADED`, required coordination/session outage는 `NOT_READY`다. - cache role eviction/OOM에서 source concurrency와 queue가 bounded다. ## Task 6 — Edge rate limit end-to-end **Files** - Modify: `src/shared-contract/src/main/java/dev/caskeleton/shared/ratelimit/*` - Modify: `src/adapter/inbound/web/src/main/java/dev/caskeleton/adapter/inbound/web/ratelimit/*` - Modify: `src/adapter/outbound/cache-redis/src/main/java/.../redis/*rate*` - Modify: `src/app-bootstrap` composition **Tests first** - inbound가 process-local map이 아니라 `EdgeRateLimitPort`를 호출한다. - subject는 raw principal/IP가 아닌 bounded pseudonymous digest다. - fixed/sliding-counter/token-bucket reference/property/concurrency vector를 통과한다. - evaluation ID replay가 quota를 두 번 소비하지 않는다. - bounded local emergency는 configured degraded provider일 때만 동작한다. - Redis/local/disabled provider exclusivity, shadow/degraded source, 429/503와 `Retry-After` mapping을 검증한다. - legacy unbounded map과 silent primary fallback을 제거한다. ## Task 7 — Idempotency v2와 Redis provider **Files** - Replace/extend `src/application-core/.../idempotency` with owner-safe v2 contracts. - Add Redis idempotency state programs/provider/codec. - Migrate the existing JPA provider to the same semantic contract only after checking its separate worktree changes; never overwrite concurrent persistence work. **Tests first** - atomic claim, fingerprint mismatch, owner/attempt-safe start/renew/complete/fail/release/inspect. - processing TTL과 replay TTL 분리. - expired `CLAIMED` takeover, expired `EXECUTING -> RECOVERY_REQUIRED`. - response-loss replay/reconciliation, conflicting response digest reject. - unverified cross-store effect는 자동 discard/re-execution하지 않는다. - JDBC/Redis provider가 같은 scope를 동시에 claim하지 않는다. **Implementation** - Redis가 cross-store exactly-once를 보장한다고 표현하지 않는다. - JPA migration 충돌이 있으면 Redis completion의 명시적 integration blocker로 보고하고 해당 worktree의 결과와 재대조한다. ## Task 8 — Efficiency lease와 optional fenced coordination **Tests first** - acquire/inspect/renew/release가 owner+operation token을 비교한다. - response loss는 `UNKNOWN/INDETERMINATE`이며 same token inspect로 reconcile한다. - expired old owner는 renew/release할 수 없다. - watchdog는 bounded scheduler와 cancellation을 사용하고 lost 상태를 전달한다. - fenced card를 선택하면 durable epoch/high-watermark 등록과 protected-resource stale-token reject를 실제 fixture로 증명한다. **Implementation** - close-only `DistributedLock`은 compatibility facade로 유지하되 새 코드가 strong lock으로 오해하지 않게 guarantee를 명명한다. - fencing 없는 Redis lease를 business correctness lock으로 광고하지 않는다. ## Task 9 — Redis Session과 JWT/session exclusive composition **Files** - Add direct `spring-session-core` and `spring-session-data-redis` to Redis leaf. - Add adapter-internal versioned session store/programs/serializer. - Add inbound web cookie/CSRF/fixation settings and security configuration. - Add app-bootstrap `jwt|redis-session` exclusive composition. **Tests first** - JWT mode는 session Redis connection/bean/thread side effect가 0이다. - pod A create/save, pod B read/touch/logout. - idle/absolute expiry, rotation, old ID reject, stale save after logout reject. - explicit allowlisted serializer N/N-1 and corrupt payload re-auth. - secure/httpOnly/SameSite/host-only cookie, CSRF enabled, fixation rotation. - repository outage/noeviction OOM/failover는 fail-open 인증으로 바뀌지 않는다. - indexed repository는 별도 opt-in이며 Cluster event cleanup 한계를 독립 검증한다. ## Task 10 — Real-service, topology, fault와 readiness Gradle tasks **Files** - Create: `src/adapter/outbound/cache-redis/src/redisTest/**` - Modify: `src/adapter/outbound/cache-redis/build.gradle` - Modify: `src/build.gradle` - Create/update Redis test topology resources and sanitized evidence reporter **Public tasks** - `redisStandaloneTest`, `redisSecurityTest`, `redisSentinelTest`, `redisClusterTest`, `redisFaultTest`, `redisCompatibilityTest` - capability card test/readiness tasks named exactly as Redis deep design §37.22 - root `redisProductionReadiness`, `redisAllImplementedCandidates` **Rules** - selected evidence에서 Docker/service 부재나 0 discovered tests는 failure다. - unselected card는 skipped가 아니라 `not selected`다. - image/program/config digest와 sanitized JUnit/topology timeline을 evidence artifact로 남긴다. ## Task 11 — Container topology와 3-node k3s qualification 이번 실행은 deep design §37.13/Phase 5A의 Sentinel-first slice만 다룬다. Cluster, fenced coordination, R3 long chaos/soak, k3s control-plane HA, physical host/AZ failure, full credential/certificate rotation은 후속 작업이다. ### Task 11.1 — Lab lifecycle contract와 host isolation RED 이 작업은 리뷰 경계를 다음처럼 분리한다. 두 하위 작업이 모두 독립 리뷰를 통과하기 전에는 부모 Task 11.1을 완료로 표시하지 않는다. - `Task 11.1A-1`: VM lifecycle, ownership marker/state, lock/signal/handoff cleanup, host fingerprint와 bounded command. 현재 구현을 동결한다. - `Task 11.1A-2`: pinned K3s generated-kubeconfig strict validator/renderer. 실행 계획은 `docs/superpowers/plans/2026-07-30-redis-lab-strict-kubeconfig-renderer.md`를 따른다. 2026-07-30 상태: `Task 11.1A-1` lifecycle/ownership과 `Task 11.1A-2` strict renderer는 whole-task 독립 review에서 Critical `0`, Important `0`, Minor `0`, SPEC PASS / QUALITY APPROVED를 받았다. fresh direct/Gradle fake-only 검증도 통과해 부모 `Task 11.1A`의 fake-only 범위는 완료다. 이는 live VM/k3s/kubectl/network/host qualification이나 Redis R2 readiness 완료를 의미하지 않는다. **Tracked files** - Create: `infra/redis-lab/README.md` - Create: `infra/redis-lab/versions.env` - Create: `infra/redis-lab/bin/redis-lab` - Create: `infra/redis-lab/cloud-init/node.yaml` - Create: `infra/redis-lab/test/redis-lab-contract.sh` - Modify: Redis Gradle VM-free lifecycle contract task **Tests first** - VM 이름은 `ca-redis-lab-server`, `ca-redis-lab-agent-1`, `ca-redis-lab-agent-2` exact allowlist만 허용한다. - server 1 + agent 2, resource `2/3GiB/12GiB`, `2/2.5GiB/12GiB`, `2/2.5GiB/12GiB`, pod CIDR `10.52.0.0/16`, service CIDR `10.53.0.0/16`, context `ca-redis-lab`을 검증한다. - host 관측은 default kubeconfig의 run-scoped copy와 원래 host context를 사용하고 read-only allowlist만 허용한다. lab 호출은 별도 ignored `src/build/redis-lab/kubeconfig`와 exact `ca-redis-lab` context를 사용한다. - default kubeconfig merge/write, host context mutation, wildcard VM cleanup, global `multipass purge`를 정적/동적 contract가 거절한다. - preflight/postflight host kubeconfig/context/node/workload fingerprint가 다르면 실패한다. - CI는 retain-on-failure를 거절하고, local opt-in만 exact VM 보존을 허용한다. - fake `multipass`/`kubectl`을 주입하는 shell contract는 partial-create cleanup과 exact command allowlist를 VM 생성 없이 검증하고 `redisLabContractTest`로 module `check`에 연결한다. - launch 전 exact name을 run-owned `PENDING`으로 atomic 예약하고 성공 직후 `CREATED`로 승격한다. timeout/실패/상태 승격 실패는 이 run이 예약한 exact name만 정리한다. - private run-scoped rendered cloud-init은 non-secret `RUN_ID|VM_NAME` ownership marker를 기록한다. cleanup/down은 bounded marker read가 state owner와 exact name 일치를 증명할 때만 delete한다. launch timeout/error는 `RECONCILE` tombstone과 bounded late-create poll로 처리하며 absent/unreadable/mismatch는 delete/state removal 없이 fail-closed한다. - lifecycle 전체는 nonblocking exclusive lock과 run identity를 사용한다. direct `up`과 `run` 모두 첫 launch 전 emergency cleanup을 활성화하고, signal/concurrent 실행이 다른 run state나 VM을 채택·삭제하지 못한다. user command에는 lock file descriptor를 상속하지 않으며 기본 bounded external child도 FD를 닫고 lock acquisition만 예외로 유지한다. `run`의 inner `up` 성공과 user command 시작 사이에도 cleanup-required flag가 연속 유지돼 zero-ownership handoff gap이 없어야 한다. - host kubeconfig copy는 fingerprint/CIDR 관측 범위가 끝나면 성공/실패와 무관하게 제거한다. - lab kubeconfig renderer는 denylist/generic-count 보강을 사용하지 않는다. pinned K3s의 canonical block-style one-cluster/context/user grammar를 별도 tracked AWK state machine으로 allowlist하며, catch-all pass-through 없이 duplicate/extra/reordered/unknown/flow-style identity와 모든 비허용 구조를 fail-closed로 거절한다. - external command와 3-node Ready 대기는 bounded이고, host service CIDR은 assigned ClusterIP에서 추측하지 않고 명시적 validated input 또는 신뢰 가능한 host 설정에서 얻는다. - mutable `curl | sudo sh` installer는 금지한다. exact K3s release URL과 SHA-256을 repository에 pin하고 host download와 각 VM transfer 뒤 다시 검증한 후에만 install/start한다. - shell contract는 별도 fixture repository에서 실행하고 actual `src/build/redis-lab` canary를 byte-for-byte 보존한다. fake PATH는 explicit safe wrapper 외 모든 명령을 fail-closed한다. ### Task 11.2A — Sentinel manifest와 security static contract GREEN **Tracked files** - Create: `infra/redis-lab/config/redis.conf.tmpl` - Create: `infra/redis-lab/config/sentinel.conf.tmpl` - Create: `infra/redis-lab/config/redis-users.acl.tmpl` - Create: `infra/redis-lab/config/sentinel-users.acl.tmpl` - Create: `infra/redis-lab/k3s/namespace.yaml` - Create: `infra/redis-lab/k3s/redis-data.yaml` - Create: `infra/redis-lab/k3s/redis-sentinel.yaml` - Create: `infra/redis-lab/k3s/network-policy.yaml` - Create: `src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/RedisLabManifestContractTest.java` - Modify: Redis Gradle manifest contract task - static contract와 live security evidence를 분리한다. YAML/템플릿 정적 통과는 TLS handshake, ACL authorization, CNI enforcement, scheduling/failover의 실행 증거가 아니다. - data Redis 3개와 Sentinel 3개는 각각 stable ordinal/headless DNS가 필요한 StatefulSet으로 구성하고 `kubernetes.io/hostname` required anti-affinity와 `maxSkew=1/DoNotSchedule` topology spread, `podManagementPolicy: Parallel`을 적용한다. - data는 PVC + AOF `appendfsync everysec`를 사용한다. Sentinel은 공식 동작상 writable config에 discovery/failover 상태를 rewrite하므로, bootstrap source를 pod별 writable PVC config로 최초 1회 atomic init-copy하고 restart 때 기존 rewritten config를 덮어쓰지 않는다. 비어 있거나 손상된 기존 config는 자동 복구로 덮지 않고 startup을 실패시킨다. - Redis image SSOT는 `src/gradle/redis-test-images.properties`의 `redis.minimum.image` exact tag+digest다. `redis.approved.image`나 임의 YAML image를 이 minimum-version Sentinel slice에 섞지 않는다. - plaintext port는 data/Sentinel 모두 0이고 TLS port만 연다. `tls-replication yes`, hostname resolution/announcement와 certificate SAN용 stable DNS를 사용한다. data plane과 Sentinel plane은 서로 다른 CA/leaf material을 가지며, peer 연결에 필요한 root만 명시적 trust bundle로 교차 포함한다. - ACL identity를 하나의 `redis-user`로 합치지 않는다. - application data user: 선택 capability/program command/key/channel만; - replica user: `+psync +replconf +ping`; - Sentinel-to-data user: 공식 최소 Sentinel control command/channel set; - Sentinel peer user: Sentinel 간 통신에 필요한 동일 superuser credential; - application Sentinel discovery user: auth/hello/ping/role과 allowlisted read-only `SENTINEL` subcommand만. default user는 off이며 application/data/discovery user에 `+@all`, `allkeys`, `allchannels`를 주지 않는다. - Redis data ACL과 Sentinel ACL은 별도 template/projection이다. Sentinel peer superuser가 data Redis에, data capability user가 Sentinel에 존재하면 static contract가 실패한다. - Secret/CA/private key/rendered config는 run별 `umask 077` 아래 생성하고 tracked manifest에는 Secret value, PEM, password가 없다. probe/command line에 `--pass`를 쓰지 않는다. - exec probe를 사용해 kubelet source CIDR 예외를 만들지 않는다. default-deny ingress/egress 뒤 data 6379, Sentinel 26379, kube-dns와 exact qualification/application pod selector만 허용한다. - Service는 headless/ClusterIP만, PDB는 data/Sentinel 각각 `minAvailable: 2`, container는 non-root, read-only root filesystem, privilege escalation false, capabilities drop ALL, seccomp RuntimeDefault, explicit requests/limits를 요구한다. - structural positive test와 한 필드씩 제거/변조한 mutation-negative fixture가 anti-affinity, spread, PDB, probes, TLS-only, ACL separation, Secret reference, NetworkPolicy, image SSOT를 실제로 fail시키는지 검증한다. - `hostPath`, `hostNetwork`, `hostPID`, `hostIPC`, privileged, NodePort, LoadBalancer, tracked Secret data/stringData/PEM과 implicit latest image를 거절한다. - static validator는 exact document inventory, duplicate YAML key/identity, selector/template 일치, exact NetworkPolicy edge graph를 검증한다. 정적 ordinal bootstrap은 최초 `redis-data-0` primary와 두 replica만 증명하며, failover 뒤 old-primary 재합류와 stale direct write 차단은 live gate에 남긴다. ### Task 11.2B — Sentinel workload와 live security baseline GREEN - Redis primary 1 + replica 2와 Sentinel 3/quorum 2를 세 node에 분산한다. - anti-affinity/topology spread, PDB, NetworkPolicy, separate data/Sentinel CA와 named ACL을 적용한다. - secret/certificate/k3s token은 매 run `umask 077` transient material로 생성하고 tracked manifest에는 값/PEM을 넣지 않는다. Sentinel bootstrap config는 Secret volume에서 pod별 writable PVC로 최초 1회 atomic init-copy하며, 기존 rewritten config를 덮어쓰지 않는다. - Redis image는 `redis.minimum.image` exact image/digest를 render하고 실제 pod image ID/digest가 일치하는지 수집한다. - data credential/CA로 Sentinel discovery가 실패하고 Sentinel material로 data command가 실패하는 negative test, untrusted CA/hostname mismatch/plaintext rejection을 실행한다. - `SENTINEL CKQUORUM`, writable config rewrite/restart, exact 3 Ready placement, PDB, default-deny/explicit-allow NetworkPolicy enforcement를 live k3s에서 검증한다. - failover 중 죽어 있던 old primary가 재합류할 때 readiness가 stale direct write를 허용하지 않고 새 primary의 replica로 수렴하는지 live 검증한다. ### Task 11.3 — Sentinel client runtime TDD - current `UnsupportedOperationException`을 먼저 고정하는 test를 quorum-consistent discovery와 분리된 discovery/data material contract로 교체한다. - 2-of-3 Sentinel이 같은 primary를 보고할 때만 후보를 만들고 loopback/wildcard/unexpected endpoint를 거절한다. - active Sentinel role이 있을 때만 registry당 daemon worker 1개, role당 fixed-delay task 1개를 만들고 `sentinel-discovery-refresh-period`(기본 30초, 5초..5분)를 적용한다. - scheduled poll과 command failure-triggered immediate rediscovery는 role별 같은 single-flight를 공유한다. `snapshot()`은 보조 trigger일 뿐 정상 polling을 대신하지 않는다. - 정상 poll은 Sentinel material만 해석하고 현재 route identity와 같으면 data material/client를 만들지 않는다. 바뀐 quorum-approved endpoint에만 data candidate를 연다. - command failure listener는 route lease 반환 뒤 topology/connectivity `UNAVAILABLE`에만 동작하며 listener 실패가 원래 certainty를 덮어쓰지 않는다. - 새 data runtime은 version/program/semantic readiness를 통과한 뒤 router에 install한다. - opaque route identity와 monotonic generation token으로 stale/same-primary candidate를 거절하고, install된 경우 old runtime은 new admission을 닫고 bounded drain/close한다. - close는 task/worker를 bounded 종료하고 late candidate를 install하지 않고 정확히 한 번 닫는다. - mutation을 자동 replay하지 않고 실행 여부가 불명확하면 `INDETERMINATE`를 보존한다. ### Task 11.4 — Multi-pod normal/failover qualification 1. host/lab preflight와 3 node/Sentinel quorum readiness를 수집한다. 2. 서로 다른 application pod에서 rate limit evaluation replay, idempotency claim/start/renew/complete, session create/read/touch/rotate/revoke를 검증한다. 3. current primary pod를 kill하고 readiness unavailable timestamp를 기록한다. 4. Sentinel quorum election, client rediscovery, runtime generation swap/drain, semantic readiness recovery를 실제 순서대로 기록한다. 5. election 60초, 추가 rediscovery/swap 30초, 총 recovery 90초의 regression limit을 적용한다. 6. rate state가 조용히 reset되지 않고 idempotency owner/terminal 결과가 중복되지 않으며 confirmed session state가 유지되는지 확인한다. 7. old primary의 replica 재합류와 모든 actor의 동일 generation 관측을 확인한다. correctness role에는 bounded `min-replicas-to-write`/`min-replicas-max-lag`와 명시적 replica acknowledgement policy를 사용한다. zero-data-loss/strong consistency를 주장하지 않으며 response-only cut 등 실행 여부가 불확실한 mutation은 `INDETERMINATE`이고 blind retry하지 않는다. ### Task 11.5 — Evidence와 exact teardown - actual image digest/image ID, config/program digest, sanitized fault/election/recovery timeline, capability별 outcome/certainty, Kubernetes/Sentinel 관측을 allowlist schema로 생성한다. - `NOT_CAPTURED` placeholder는 qualification 성공으로 인정하지 않는다. - sanitizer/reconciler 성공 뒤에도 human clean commit/remote CI 전에는 `releaseQualification=NOT_CLAIMED`를 유지한다. - 성공/실패 모두 exact VM allowlist를 teardown하고 lab resource가 0인지 확인한다. local retain-on-failure opt-in은 명시된 경우만 허용하고 CI에서는 금지한다. ## Task 12 — CI, runbook, verification와 Wiki capture **CI** - PR blocking `redis-standalone` job을 `release-gate.needs`와 result loop에 실제 포함한다. - nightly/release Redis production readiness workflow를 추가한다. - workflow contract test로 blocking job/aggregator 집합 동등성을 검증한다. **Verification** ```bash cd src ./gradlew :application-core:redisPolicyContractTest --console=plain ./gradlew :shared-contract:edgeRateLimitContractTest --console=plain ./gradlew :adapter:outbound:cache-redis:check --console=plain ./gradlew :app-bootstrap:redisCompositionTest --console=plain ./gradlew redisProductionReadiness --console=plain ./gradlew test --console=plain ./gradlew check --console=plain ./gradlew verifyCleanArchitectureDependencies --console=plain ./gradlew verifyPublicPathSnapshot --console=plain ./gradlew verifyEnvKeys --console=plain ``` **Documentation** - capability별 실제 readiness와 남은 R3 한계를 README/spec/runbook에 동기화한다. - 실행 명령, image/config/program digest, 실패/차단을 public LLM Wiki `/home/donghyeon/workspace/ai-tools/llm-wiki/raw/branch-notes/main.md`에 기록하고 실제 파생 오류/면접/블로그 raw 문서를 양방향 링크한다. **Completion gate** - Task 11의 exit gate를 통과하면 `Sentinel-first R2-ready candidate`라고만 보고한다. - clean committed source와 실제 remote CI가 없으면 selected/R2로 승격하지 않는다. - 이 milestone 보고 뒤 멈추고 Cluster/R3/fenced coordination 또는 fileserver/HTTP client 중 다음 작업을 사용자와 다시 정한다. ## Task 13 — Resume blocker: selection-driven role activation과 default boot **Problem** - provider definition뿐 아니라 role binding도 capability가 선택되지 않으면 inert여야 한다. - 현재 구현은 role binding 전체를 runtime으로 열고 health contributor도 role property 존재만으로 활성화한다. - local 기본값에서 inbound rate-limit은 provider 없이 활성화되면 안 된다. **Tests first** - CACHE/COORDINATION/SESSION deployment와 role을 모두 사전 선언해도 cache/rate/idempotency/lease/ session capability가 비활성이면 credential/trust resolution, native client, scheduler/subscriber, Redis health contributor가 모두 0이다. - 각 capability가 `redis`를 선택할 때만 해당 role이 활성화된다. - 같은 role을 쓰는 coordination capability 둘 이상은 하나의 runtime만 공유한다. - 선택 capability의 role binding이 빠지면 material resolution 전에 startup이 실패한다. - shipped `.env`와 실제 `application.yml`은 transport disabled/provider disabled 조합으로 기동 가능하고 중복 legacy rate-limit block이 없다. **Implementation** - deployment/role registry validation과 runtime activation을 분리한다. - `selectedCapabilities`가 비어 있는 role은 registry/router/health에서 제외한다. - bootstrap health condition도 role property가 아니라 effective selected capability로 판단한다. - provider 설정은 inert 후보로 남기되 선택된 capability의 잘못된 role은 fail closed 한다. ## Task 14 — Resume blocker: capability-aware semantic readiness **Problem** - PING만으로 `AVAILABLE/PROBE_SUCCEEDED`를 선언하지 않는다. - required coordination/session은 실제 선택 capability의 program ACL과 최소 read/write 계약이 동작해야 ready다. **Tests first** - PING은 성공하지만 `SCRIPT LOAD`/`EVALSHA`가 ACL로 거절된 coordination/session user는 `redisRequired=DOWN`이다. - capability별 representative program의 실제 key count와 command-to-key mapping을 그대로 검증한다. rate-limit의 state/dedup/order key와 session tombstone key 중 하나만 ACL pattern에서 빠져도 semantic readiness는 실패한다. - Redis 7.2 미만 server는 metadata 표기만으로 통과하지 않고 bounded runtime handshake에서 sanitized unsupported-version 상태가 된다. - 대표 program과 ACL probe script가 이미 warm인 상태에서도 runtime user의 `SCRIPT LOAD` 권한 누락을 별도로 탐지한다. - cache optional role에서 semantic probe 실패는 application liveness/readiness를 내리지 않고 `DEGRADED`만 보고한다. - 선언된 optional cache가 cold-start connect/PING에 일시 실패해도 context는 bounded unavailable route로 시작하고, health-triggered bounded single-flight reconnect 뒤 재시작 없이 복구한다. invalid configuration/material/program/schema는 계속 startup failure이며 required coordination/session은 fail closed다. - probe는 raw key/value, credential, server exception을 health detail에 노출하지 않는다. - probe key는 bounded, namespaced, TTL이 있고 성공/실패 후 잔여 상태가 없다. - saturation/recent command failure/closed route를 distinct sanitized reason으로 분류한다. - health scrape는 role별 minimum cadence와 single-flight로 full semantic suite 실행을 제한하고, cached observation의 시각/age를 노출해 stale success를 숨기지 않는다. **Implementation** - role별 선택 capability를 입력으로 immutable semantic probe plan을 만든다. - probe는 catalog-owned bounded program과 capability-safe ephemeral operation만 사용한다. - optional cold-start outage는 resource-free unavailable runtime과 bounded on-demand reconnect로 표현하며 별도 unbounded scheduler/thread를 만들지 않는다. L1 invalidation subscription은 route recovery 시 실제 runtime에 다시 연결된다. - eviction은 runtime `CONFIG` 권한을 열지 않고 `CONFIGURED_EXPECTATION_ONLY`로 유지하며 외부 attestation 미완료를 readiness detail에 명시한다. ## Task 15 — Resume blocker: bounded common primitive catalog **Problem** - Deep design §14.6–§14.9의 자주 쓰는 race-safe helper가 아직 compare/delete 중심 R0 foundation에 머물러 있다. **Tests first** - String, counter, hash, set, sorted-set, list baseline은 typed/versioned key, value/count/byte/deadline, role, slot, TTL, certainty bound를 강제한다. - bitmap/HLL/geo는 billing/auth correctness에 사용할 수 없는 explicit semantic classification과 offset/result/fan-in bound를 강제한다. - `INCR -> EXPIRE`, set/list admission, revision-CAS는 실제 Redis concurrency에서 atomic하다. - unbounded `HGETALL`, `SMEMBERS`, `LRANGE`, arbitrary command/script surface는 제공하지 않는다. **Implementation** - package-private `RedisPrimitiveCatalog`과 structure별 bounded facade를 Redis leaf 내부에 둔다. - application/shared public API에는 Redis command나 raw key를 노출하지 않는다. - 아직 실제 semantic consumer가 없는 primitive는 Spring bean/public capability로 노출하지 않는다. ## Task 16 — Resume blocker: capability observability와 graceful lifecycle **Tests first** - cache/rate/idempotency/lease/session의 operation, outcome, certainty, role, queue/latency가 bounded low-cardinality metric/event로 관측된다. - raw key, subject, session/idempotency/lease token, secret reference/value, exception message는 tag/log/trace에 들어가지 않는다. - optional cache와 required coordination/session의 failure signal이 health와 metric에서 일치한다. - shutdown은 subscriber/scheduler/router/runtime 순서로 bounded drain되고 새 command를 거절한다. **Implementation** - framework-neutral observation event/port와 Micrometer rendering을 계층 소유권에 맞게 둔다. - trace/log는 기존 skeleton observability 경계를 재사용하고 Redis native type을 core에 유출하지 않는다. - `docs/registries/metrics.yaml`과 runbook을 실제 emitted metric과 동기화한다. ## Task 17 — Resume final review, readiness truth, verification와 Wiki - Task 13–16을 task별 spec/code-quality review한다. - Redis deep design §39/§40을 독립 재검토해 selected/implemented-candidate/not-implemented를 실제 evidence와 일치시킨다. - Sentinel/Cluster/k3s/R3 evidence가 없으면 지원/완료로 표기하지 않는다. - Task 12의 전체 검증을 실행하고 동시 작업의 비-Redis 실패는 소유 파일과 증거를 분리한다. - Redis README/spec/runbook, readiness registry, CI artifact 계약을 동기화한다. - LLM Wiki branch-note와 실제 파생 raw 문서를 양방향 링크로 캡처한다.