Files
clean-architecture-backend-…/docs/superpowers/plans/2026-07-29-redis-production-capability-completion.md
T

37 KiB
Raw Blame History

Redis Production Capability Completion Plan

Scope: Redis를 먼저 완료한다. 현재 실행 단위는 deep design Phase 5 전체가 아니라 Sentinel-first R2 qualification slice다. 이 slice의 검증과 보고가 끝나면 멈추고 fileserver, HTTP client, Redis Cluster/R3 중 다음 우선순위를 다시 정한다.

Workflow note: 저장소가 지정한 Superpowers 설계·계획·TDD·디버깅·검증·리뷰 워크플로우를 적용한다. agent는 human-only commit 정책에 따라 stage/commit/amend/push하지 않는다.

Goal: 2026-07-26-redis-production-capability-design.md의 Phase 15를 capability별로 구현하고, standalone 기능의 존재를 production readiness로 오표기하지 않는 Redis platform을 만든다.

Architecture: application-coreshared-contract는 provider-neutral semantic contract만 소유한다. adapter:outbound:cache-redis가 Redis deployment, topology, key, codec, program, runtime과 capability provider를 소유한다. adapter:inbound:web은 HTTP rate/session 보안 매핑만, app-bootstrap은 provider/role/auth-mode composition만 소유한다. domain-core에는 Redis 개념을 추가하지 않는다.

Readiness rule: Redis leaf 전체에 단일 R2 label을 부여하지 않는다. redis-cache, redis-edge-rate-limit, redis-request-replay-idempotency, redis-cache-refresh-soft-lease, redis-fenced-coordination, redis-session card가 독립적으로 승격한다. R3 증거가 없는 failover/reshard/rotation은 R2 범위로 과장하지 않는다.

Worktree rule: 현재 main worktree의 다른 기술 변경은 사용자 소유다. Redis가 소유하지 않는 fileserver, HTTP client, messaging, notification, object storage 변경을 되돌리거나 포맷하지 않는다.

Current milestone exit: agent-side 목표는 R2-ready candidate다. clean committed source와 실제 remote GitHub Actions evidence가 없으면 card를 selected로 바꾸거나 R2라고 주장하지 않는다.


Task 0 — Baseline과 acceptance registry 고정

Files

  • Create: src/config/redis/readiness-cards.yaml
  • Create: src/gradle/redis-test-images.properties
  • Modify: src/adapter/outbound/cache-redis/README.md
  • Modify: docs/superpowers/specs/2026-07-26-redis-production-capability-design.md

Tests first

  • registry가 canonical card ID 여섯 개를 정확히 한 번 포함하는지 실패 테스트를 작성한다.
  • image tag에 exact version과 digest가 없으면 configuration이 실패하는 테스트를 작성한다.
  • selected, implemented-candidate, not-implemented 이외 상태를 거절한다.
  • 현재 구현과 다른 readiness 표기를 거절한다.

Implementation

  • 시작 상태는 cache/rate를 implemented-candidate, 나머지는 not-implemented로 기록한다.
  • 실제 required evidence가 생기기 전에는 어떤 card도 selected R2로 승격하지 않는다.
  • Redis minimum version은 실행 가능한 image/digest와 program manifest를 한 SSOT로 맞춘다.

Verification

cd src
./gradlew :adapter:outbound:cache-redis:test --tests '*RedisReadinessRegistryTest' --console=plain

Task 1 — Canonical deployment/topology/role model

Files

  • Create: src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderProperties.java
  • Create: src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettings.java
  • Create: src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactory.java
  • Create: src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRole.java
  • Create: src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRoleBinding.java
  • Test: src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactoryTest.java
  • Test: src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderPropertiesBindingTest.java

Tests first

  • topology는 standalone|sentinel|cluster 중 정확히 하나다.
  • endpoint는 non-empty, unique, bounded host/port다.
  • Sentinel은 master name, 최소 3개 discovery endpoint, data/Sentinel auth와 TLS를 분리한다.
  • Cluster는 database 0만 허용하고 seed가 비어 있으면 실패한다.
  • role은 존재하는 deployment만 참조한다.
  • cache와 session/coordination의 incompatible co-location을 startup 전에 거절한다.
  • provider 정의만 있고 capability binding이 없으면 runtime side effect가 0이다.

Implementation

  • Spring binding class와 validated sealed runtime model을 분리한다.
  • legacy app.cache.redisapp.rate-limit은 migration compiler 입력으로만 허용하고 canonical model과 동시에 설정되면 precedence를 정하지 않고 실패한다.
  • ClientMode.EXTERNAL을 topology로 취급하지 않는다.

Verification

cd src
./gradlew :adapter:outbound:cache-redis:test --tests '*RedisDeploymentSettings*' --console=plain

Task 2 — Topology-aware runtime, TLS/ACL과 secret material

Files

  • Create: .../redis/runtime/RedisDeploymentRuntime.java
  • Create: .../redis/runtime/RedisDeploymentRuntimeFactory.java
  • Create: .../redis/runtime/StandaloneRedisDeploymentRuntime.java
  • Create: .../redis/runtime/SentinelRedisDeploymentRuntime.java
  • Create: .../redis/runtime/ClusterRedisDeploymentRuntime.java
  • Create: .../redis/security/RedisCredentialMaterialProvider.java
  • Create: .../redis/security/RedisCredentialRotationCoordinator.java
  • Modify: src/adapter/outbound/cache-redis/build.gradle
  • Modify: src/adapter/outbound/cache-redis/gradle.lockfile

Tests first

  • standalone/Sentinel/Cluster가 각자 다른 native client/runtime을 만든다.
  • Sentinel discovery credential/trust와 data-node credential/trust가 섞이지 않는다.
  • Cluster client는 periodic+adaptive topology refresh, DB 0, bounded redirect/queue profile을 가진다.
  • production profile에서 plaintext, trust-all, hostname verification off를 거절한다.
  • named ACL username이 없거나 raw password가 YAML에 있으면 production activation이 실패한다.
  • duplicate/out-of-order rotation event, expiry 재조회, new connection 검증 실패가 old traffic을 안전하게 보존한다.
  • disabled capability는 client/event-loop/subscriber/scheduler를 만들지 않는다.

Implementation

  • direct spring-data-redis, lettuce-core dependency를 leaf가 소유한다.
  • deployment별 client resources와 lifecycle을 소유한다.
  • connect/TLS/acquire/command/overall/shutdown timeout을 분리한다.
  • 기존 no-replay, disconnected reject, finite queue/count/byte admission을 topology runtime에도 보존한다.
  • secret value/reference/provider exception을 log/metric에 남기지 않는다.

Task 3 — Key, codec, program manifest foundation

Files

  • Create: src/config/redis/program-set.schema.json
  • Modify: src/adapter/outbound/cache-redis/src/main/resources/redis/program-set.json
  • Modify: src/adapter/outbound/cache-redis/src/main/resources/redis/rate-program-set.json
  • Modify: .../redis/RedisProgramDescriptor.java
  • Modify: .../redis/RedisProgramCatalog.java
  • Modify: .../redis/RedisLuaProgramExecutor.java
  • Create: .../redis/key/RedisKeyMaterialProvider.java
  • Create: .../redis/codec/RedisCapabilityCodec.java

Tests first

  • 모든 program은 exact source digest, semantic version, ordered KEYS/ARGV, result schema, slot rule, state/TTL bound, minimum Redis version, retry/certainty, ACL command를 가진다.
  • manifest와 Java descriptor가 drift하면 build가 실패한다.
  • NOSCRIPT recovery는 bounded SCRIPT LOAD -> EVALSHA이고 arbitrary source 실행 surface가 없다.
  • same-resource multi-key는 real CLUSTER KEYSLOT과 같은 slot이다.
  • key digest material rotation은 fixed/dual-read-delete/cold-cutover rule을 지킨다.
  • cache/idempotency/session codec은 N/N-1, future/corrupt/oversize/forbidden type을 구분한다.

Implementation

  • foundation/rate manifest를 하나의 versioned registry contract로 통합하되 capability package와 facade는 분리한다.
  • raw command, raw key, generic program executor를 Spring/application public surface에 노출하지 않는다.

Task 4 — Cache consistency spine와 semantic region composition

Files

  • Modify: src/application-core/src/main/java/dev/caskeleton/application/cache/*
  • Create: .../redis/cache/RedisCacheGenerationStore.java
  • Create: .../redis/cache/RedisCacheRegionCompiler.java
  • Add resources: region-generation-init-v1.lua, region-generation-bump-v1.lua, cache-record-if-generation-v1.lua
  • Modify: .../redis/RedisStringCacheRegion.java
  • Tests: application barrier tests, Redis real-service concurrency tests, binding tests

Tests first

  • source load 중 generation bump가 일어나면 old result가 visible하지 않다.
  • captured generation과 source revision이 바뀌면 stale writer가 새 값을 덮어쓰지 않는다.
  • generation init race에서 하나의 canonical generation만 선택된다.
  • operation ID가 같은 bump replay는 한 번만 적용된다.
  • 여러 semantic region의 duplicate/missing binding은 fail-fast다.
  • 실제 consumer가 semantic CacheRegionPortCacheAsideExecutor를 사용하고 legacy fail-open router와 암묵적으로 섞이지 않는다.

Implementation decision

  • source revision은 opaque하므로 lexical “newer” 비교를 하지 않는다.
  • region generation은 mass invalidation fence다.
  • per-key invalidation은 해당 key의 revision/tombstone fence를 사용해 region 전체를 bump하지 않는다.
  • write는 captured generation/revision condition을 만족할 때만 기록한다.

Task 5 — Distributed refresh soft lease, L1/L2와 cache observability

Files

  • Create application cache refresh coordination contracts without Redis types.
  • Create Redis refresh claim/release programs and semantic provider.
  • Create bounded L1 cache decorator and invalidation subscriber/reconciler.
  • Create framework-free cache observation events and Micrometer adapter instrumentation.
  • Update docs/registries/metrics.yaml.

Tests first

  • 두 pod simulation에서 정상 시 refresh owner는 하나다.
  • lease expiry에서는 duplicate load를 허용하지만 generation guard가 stale write를 차단한다.
  • disconnected invalidation subscriber는 L1을 flush하고 generation을 재확인한다.
  • Pub/Sub event loss에도 L1 TTL/generation reconciliation으로 stale bound를 지킨다.
  • L1 max weight/cardinality/TTL, subscriber queue, refresh scheduler가 모두 bounded다.
  • Redis liveness는 애플리케이션 liveness를 내리지 않는다.
  • optional cache outage는 DEGRADED, required coordination/session outage는 NOT_READY다.
  • cache role eviction/OOM에서 source concurrency와 queue가 bounded다.

Task 6 — Edge rate limit end-to-end

Files

  • Modify: src/shared-contract/src/main/java/dev/caskeleton/shared/ratelimit/*
  • Modify: src/adapter/inbound/web/src/main/java/dev/caskeleton/adapter/inbound/web/ratelimit/*
  • Modify: src/adapter/outbound/cache-redis/src/main/java/.../redis/*rate*
  • Modify: src/app-bootstrap composition

Tests first

  • inbound가 process-local map이 아니라 EdgeRateLimitPort를 호출한다.
  • subject는 raw principal/IP가 아닌 bounded pseudonymous digest다.
  • fixed/sliding-counter/token-bucket reference/property/concurrency vector를 통과한다.
  • evaluation ID replay가 quota를 두 번 소비하지 않는다.
  • bounded local emergency는 configured degraded provider일 때만 동작한다.
  • Redis/local/disabled provider exclusivity, shadow/degraded source, 429/503와 Retry-After mapping을 검증한다.
  • legacy unbounded map과 silent primary fallback을 제거한다.

Task 7 — Idempotency v2와 Redis provider

Files

  • Replace/extend src/application-core/.../idempotency with owner-safe v2 contracts.
  • Add Redis idempotency state programs/provider/codec.
  • Migrate the existing JPA provider to the same semantic contract only after checking its separate worktree changes; never overwrite concurrent persistence work.

Tests first

  • atomic claim, fingerprint mismatch, owner/attempt-safe start/renew/complete/fail/release/inspect.
  • processing TTL과 replay TTL 분리.
  • expired CLAIMED takeover, expired EXECUTING -> RECOVERY_REQUIRED.
  • response-loss replay/reconciliation, conflicting response digest reject.
  • unverified cross-store effect는 자동 discard/re-execution하지 않는다.
  • JDBC/Redis provider가 같은 scope를 동시에 claim하지 않는다.

Implementation

  • Redis가 cross-store exactly-once를 보장한다고 표현하지 않는다.
  • JPA migration 충돌이 있으면 Redis completion의 명시적 integration blocker로 보고하고 해당 worktree의 결과와 재대조한다.

Task 8 — Efficiency lease와 optional fenced coordination

Tests first

  • acquire/inspect/renew/release가 owner+operation token을 비교한다.
  • response loss는 UNKNOWN/INDETERMINATE이며 same token inspect로 reconcile한다.
  • expired old owner는 renew/release할 수 없다.
  • watchdog는 bounded scheduler와 cancellation을 사용하고 lost 상태를 전달한다.
  • fenced card를 선택하면 durable epoch/high-watermark 등록과 protected-resource stale-token reject를 실제 fixture로 증명한다.

Implementation

  • close-only DistributedLock은 compatibility facade로 유지하되 새 코드가 strong lock으로 오해하지 않게 guarantee를 명명한다.
  • fencing 없는 Redis lease를 business correctness lock으로 광고하지 않는다.

Task 9 — Redis Session과 JWT/session exclusive composition

Files

  • Add direct spring-session-core and spring-session-data-redis to Redis leaf.
  • Add adapter-internal versioned session store/programs/serializer.
  • Add inbound web cookie/CSRF/fixation settings and security configuration.
  • Add app-bootstrap jwt|redis-session exclusive composition.

Tests first

  • JWT mode는 session Redis connection/bean/thread side effect가 0이다.
  • pod A create/save, pod B read/touch/logout.
  • idle/absolute expiry, rotation, old ID reject, stale save after logout reject.
  • explicit allowlisted serializer N/N-1 and corrupt payload re-auth.
  • secure/httpOnly/SameSite/host-only cookie, CSRF enabled, fixation rotation.
  • repository outage/noeviction OOM/failover는 fail-open 인증으로 바뀌지 않는다.
  • indexed repository는 별도 opt-in이며 Cluster event cleanup 한계를 독립 검증한다.

Task 10 — Real-service, topology, fault와 readiness Gradle tasks

Files

  • Create: src/adapter/outbound/cache-redis/src/redisTest/**
  • Modify: src/adapter/outbound/cache-redis/build.gradle
  • Modify: src/build.gradle
  • Create/update Redis test topology resources and sanitized evidence reporter

Public tasks

  • redisStandaloneTest, redisSecurityTest, redisSentinelTest, redisClusterTest, redisFaultTest, redisCompatibilityTest
  • capability card test/readiness tasks named exactly as Redis deep design §37.22
  • root redisProductionReadiness, redisAllImplementedCandidates

Rules

  • selected evidence에서 Docker/service 부재나 0 discovered tests는 failure다.
  • unselected card는 skipped가 아니라 not selected다.
  • image/program/config digest와 sanitized JUnit/topology timeline을 evidence artifact로 남긴다.

Task 11 — Container topology와 3-node k3s qualification

이번 실행은 deep design §37.13/Phase 5A의 Sentinel-first slice만 다룬다. Cluster, fenced coordination, R3 long chaos/soak, k3s control-plane HA, physical host/AZ failure, full credential/certificate rotation은 후속 작업이다.

Task 11.1 — Lab lifecycle contract와 host isolation RED

이 작업은 리뷰 경계를 다음처럼 분리한다. 두 하위 작업이 모두 독립 리뷰를 통과하기 전에는 부모 Task 11.1을 완료로 표시하지 않는다.

  • Task 11.1A-1: VM lifecycle, ownership marker/state, lock/signal/handoff cleanup, host fingerprint와 bounded command. 현재 구현을 동결한다.
  • Task 11.1A-2: pinned K3s generated-kubeconfig strict validator/renderer. 실행 계획은 docs/superpowers/plans/2026-07-30-redis-lab-strict-kubeconfig-renderer.md를 따른다.

2026-07-30 상태: Task 11.1A-1 lifecycle/ownership과 Task 11.1A-2 strict renderer는 whole-task 독립 review에서 Critical 0, Important 0, Minor 0, SPEC PASS / QUALITY APPROVED를 받았다. fresh direct/Gradle fake-only 검증도 통과해 부모 Task 11.1A의 fake-only 범위는 완료다. 이는 live VM/k3s/kubectl/network/host qualification이나 Redis R2 readiness 완료를 의미하지 않는다.

Tracked files

  • Create: infra/redis-lab/README.md
  • Create: infra/redis-lab/versions.env
  • Create: infra/redis-lab/bin/redis-lab
  • Create: infra/redis-lab/cloud-init/node.yaml
  • Create: infra/redis-lab/test/redis-lab-contract.sh
  • Modify: Redis Gradle VM-free lifecycle contract task

Tests first

  • VM 이름은 ca-redis-lab-server, ca-redis-lab-agent-1, ca-redis-lab-agent-2 exact allowlist만 허용한다.
  • server 1 + agent 2, resource 2/3GiB/12GiB, 2/2.5GiB/12GiB, 2/2.5GiB/12GiB, pod CIDR 10.52.0.0/16, service CIDR 10.53.0.0/16, context ca-redis-lab을 검증한다.
  • host 관측은 default kubeconfig의 run-scoped copy와 원래 host context를 사용하고 read-only allowlist만 허용한다. lab 호출은 별도 ignored src/build/redis-lab/kubeconfig와 exact ca-redis-lab context를 사용한다.
  • default kubeconfig merge/write, host context mutation, wildcard VM cleanup, global multipass purge를 정적/동적 contract가 거절한다.
  • preflight/postflight host kubeconfig/context/node/workload fingerprint가 다르면 실패한다.
  • CI는 retain-on-failure를 거절하고, local opt-in만 exact VM 보존을 허용한다.
  • fake multipass/kubectl을 주입하는 shell contract는 partial-create cleanup과 exact command allowlist를 VM 생성 없이 검증하고 redisLabContractTest로 module check에 연결한다.
  • launch 전 exact name을 run-owned PENDING으로 atomic 예약하고 성공 직후 CREATED로 승격한다. timeout/실패/상태 승격 실패는 이 run이 예약한 exact name만 정리한다.
  • private run-scoped rendered cloud-init은 non-secret RUN_ID|VM_NAME ownership marker를 기록한다. cleanup/down은 bounded marker read가 state owner와 exact name 일치를 증명할 때만 delete한다. launch timeout/error는 RECONCILE tombstone과 bounded late-create poll로 처리하며 absent/unreadable/mismatch는 delete/state removal 없이 fail-closed한다.
  • lifecycle 전체는 nonblocking exclusive lock과 run identity를 사용한다. direct uprun 모두 첫 launch 전 emergency cleanup을 활성화하고, signal/concurrent 실행이 다른 run state나 VM을 채택·삭제하지 못한다. user command에는 lock file descriptor를 상속하지 않으며 기본 bounded external child도 FD를 닫고 lock acquisition만 예외로 유지한다. run의 inner up 성공과 user command 시작 사이에도 cleanup-required flag가 연속 유지돼 zero-ownership handoff gap이 없어야 한다.
  • host kubeconfig copy는 fingerprint/CIDR 관측 범위가 끝나면 성공/실패와 무관하게 제거한다.
  • lab kubeconfig renderer는 denylist/generic-count 보강을 사용하지 않는다. pinned K3s의 canonical block-style one-cluster/context/user grammar를 별도 tracked AWK state machine으로 allowlist하며, catch-all pass-through 없이 duplicate/extra/reordered/unknown/flow-style identity와 모든 비허용 구조를 fail-closed로 거절한다.
  • external command와 3-node Ready 대기는 bounded이고, host service CIDR은 assigned ClusterIP에서 추측하지 않고 명시적 validated input 또는 신뢰 가능한 host 설정에서 얻는다.
  • mutable curl | sudo sh installer는 금지한다. exact K3s release URL과 SHA-256을 repository에 pin하고 host download와 각 VM transfer 뒤 다시 검증한 후에만 install/start한다.
  • shell contract는 별도 fixture repository에서 실행하고 actual src/build/redis-lab canary를 byte-for-byte 보존한다. fake PATH는 explicit safe wrapper 외 모든 명령을 fail-closed한다.

Task 11.2A — Sentinel manifest와 security static contract GREEN

Tracked files

  • Create: infra/redis-lab/config/redis.conf.tmpl

  • Create: infra/redis-lab/config/sentinel.conf.tmpl

  • Create: infra/redis-lab/config/redis-users.acl.tmpl

  • Create: infra/redis-lab/config/sentinel-users.acl.tmpl

  • Create: infra/redis-lab/k3s/namespace.yaml

  • Create: infra/redis-lab/k3s/redis-data.yaml

  • Create: infra/redis-lab/k3s/redis-sentinel.yaml

  • Create: infra/redis-lab/k3s/network-policy.yaml

  • Create: src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/RedisLabManifestContractTest.java

  • Modify: Redis Gradle manifest contract task

  • static contract와 live security evidence를 분리한다. YAML/템플릿 정적 통과는 TLS handshake, ACL authorization, CNI enforcement, scheduling/failover의 실행 증거가 아니다.

  • data Redis 3개와 Sentinel 3개는 각각 stable ordinal/headless DNS가 필요한 StatefulSet으로 구성하고 kubernetes.io/hostname required anti-affinity와 maxSkew=1/DoNotSchedule topology spread, podManagementPolicy: Parallel을 적용한다.

  • data는 PVC + AOF appendfsync everysec를 사용한다. Sentinel은 공식 동작상 writable config에 discovery/failover 상태를 rewrite하므로, bootstrap source를 pod별 writable PVC config로 최초 1회 atomic init-copy하고 restart 때 기존 rewritten config를 덮어쓰지 않는다. 비어 있거나 손상된 기존 config는 자동 복구로 덮지 않고 startup을 실패시킨다.

  • Redis image SSOT는 src/gradle/redis-test-images.propertiesredis.minimum.image exact tag+digest다. redis.approved.image나 임의 YAML image를 이 minimum-version Sentinel slice에 섞지 않는다.

  • plaintext port는 data/Sentinel 모두 0이고 TLS port만 연다. tls-replication yes, hostname resolution/announcement와 certificate SAN용 stable DNS를 사용한다. data plane과 Sentinel plane은 서로 다른 CA/leaf material을 가지며, peer 연결에 필요한 root만 명시적 trust bundle로 교차 포함한다.

  • ACL identity를 하나의 redis-user로 합치지 않는다.

    • application data user: 선택 capability/program command/key/channel만;
    • replica user: +psync +replconf +ping;
    • Sentinel-to-data user: 공식 최소 Sentinel control command/channel set;
    • Sentinel peer user: Sentinel 간 통신에 필요한 동일 superuser credential;
    • application Sentinel discovery user: auth/hello/ping/role과 allowlisted read-only SENTINEL subcommand만. default user는 off이며 application/data/discovery user에 +@all, allkeys, allchannels를 주지 않는다.
  • Redis data ACL과 Sentinel ACL은 별도 template/projection이다. Sentinel peer superuser가 data Redis에, data capability user가 Sentinel에 존재하면 static contract가 실패한다.

  • Secret/CA/private key/rendered config는 run별 umask 077 아래 생성하고 tracked manifest에는 Secret value, PEM, password가 없다. probe/command line에 --pass를 쓰지 않는다.

  • exec probe를 사용해 kubelet source CIDR 예외를 만들지 않는다. default-deny ingress/egress 뒤 data 6379, Sentinel 26379, kube-dns와 exact qualification/application pod selector만 허용한다.

  • Service는 headless/ClusterIP만, PDB는 data/Sentinel 각각 minAvailable: 2, container는 non-root, read-only root filesystem, privilege escalation false, capabilities drop ALL, seccomp RuntimeDefault, explicit requests/limits를 요구한다.

  • structural positive test와 한 필드씩 제거/변조한 mutation-negative fixture가 anti-affinity, spread, PDB, probes, TLS-only, ACL separation, Secret reference, NetworkPolicy, image SSOT를 실제로 fail시키는지 검증한다.

  • hostPath, hostNetwork, hostPID, hostIPC, privileged, NodePort, LoadBalancer, tracked Secret data/stringData/PEM과 implicit latest image를 거절한다.

  • static validator는 exact document inventory, duplicate YAML key/identity, selector/template 일치, exact NetworkPolicy edge graph를 검증한다. 정적 ordinal bootstrap은 최초 redis-data-0 primary와 두 replica만 증명하며, failover 뒤 old-primary 재합류와 stale direct write 차단은 live gate에 남긴다.

Task 11.2B — Sentinel workload와 live security baseline GREEN

  • Redis primary 1 + replica 2와 Sentinel 3/quorum 2를 세 node에 분산한다.
  • anti-affinity/topology spread, PDB, NetworkPolicy, separate data/Sentinel CA와 named ACL을 적용한다.
  • secret/certificate/k3s token은 매 run umask 077 transient material로 생성하고 tracked manifest에는 값/PEM을 넣지 않는다. Sentinel bootstrap config는 Secret volume에서 pod별 writable PVC로 최초 1회 atomic init-copy하며, 기존 rewritten config를 덮어쓰지 않는다.
  • Redis image는 redis.minimum.image exact image/digest를 render하고 실제 pod image ID/digest가 일치하는지 수집한다.
  • data credential/CA로 Sentinel discovery가 실패하고 Sentinel material로 data command가 실패하는 negative test, untrusted CA/hostname mismatch/plaintext rejection을 실행한다.
  • SENTINEL CKQUORUM, writable config rewrite/restart, exact 3 Ready placement, PDB, default-deny/explicit-allow NetworkPolicy enforcement를 live k3s에서 검증한다.
  • failover 중 죽어 있던 old primary가 재합류할 때 readiness가 stale direct write를 허용하지 않고 새 primary의 replica로 수렴하는지 live 검증한다.

Task 11.3 — Sentinel client runtime TDD

  • current UnsupportedOperationException을 먼저 고정하는 test를 quorum-consistent discovery와 분리된 discovery/data material contract로 교체한다.
  • 2-of-3 Sentinel이 같은 primary를 보고할 때만 후보를 만들고 loopback/wildcard/unexpected endpoint를 거절한다.
  • active Sentinel role이 있을 때만 registry당 daemon worker 1개, role당 fixed-delay task 1개를 만들고 sentinel-discovery-refresh-period(기본 30초, 5초..5분)를 적용한다.
  • scheduled poll과 command failure-triggered immediate rediscovery는 role별 같은 single-flight를 공유한다. snapshot()은 보조 trigger일 뿐 정상 polling을 대신하지 않는다.
  • 정상 poll은 Sentinel material만 해석하고 현재 route identity와 같으면 data material/client를 만들지 않는다. 바뀐 quorum-approved endpoint에만 data candidate를 연다.
  • command failure listener는 route lease 반환 뒤 topology/connectivity UNAVAILABLE에만 동작하며 listener 실패가 원래 certainty를 덮어쓰지 않는다.
  • 새 data runtime은 version/program/semantic readiness를 통과한 뒤 router에 install한다.
  • opaque route identity와 monotonic generation token으로 stale/same-primary candidate를 거절하고, install된 경우 old runtime은 new admission을 닫고 bounded drain/close한다.
  • close는 task/worker를 bounded 종료하고 late candidate를 install하지 않고 정확히 한 번 닫는다.
  • mutation을 자동 replay하지 않고 실행 여부가 불명확하면 INDETERMINATE를 보존한다.

Task 11.4 — Multi-pod normal/failover qualification

  1. host/lab preflight와 3 node/Sentinel quorum readiness를 수집한다.
  2. 서로 다른 application pod에서 rate limit evaluation replay, idempotency claim/start/renew/complete, session create/read/touch/rotate/revoke를 검증한다.
  3. current primary pod를 kill하고 readiness unavailable timestamp를 기록한다.
  4. Sentinel quorum election, client rediscovery, runtime generation swap/drain, semantic readiness recovery를 실제 순서대로 기록한다.
  5. election 60초, 추가 rediscovery/swap 30초, 총 recovery 90초의 regression limit을 적용한다.
  6. rate state가 조용히 reset되지 않고 idempotency owner/terminal 결과가 중복되지 않으며 confirmed session state가 유지되는지 확인한다.
  7. old primary의 replica 재합류와 모든 actor의 동일 generation 관측을 확인한다.

correctness role에는 bounded min-replicas-to-write/min-replicas-max-lag와 명시적 replica acknowledgement policy를 사용한다. zero-data-loss/strong consistency를 주장하지 않으며 response-only cut 등 실행 여부가 불확실한 mutation은 INDETERMINATE이고 blind retry하지 않는다.

Task 11.5 — Evidence와 exact teardown

  • actual image digest/image ID, config/program digest, sanitized fault/election/recovery timeline, capability별 outcome/certainty, Kubernetes/Sentinel 관측을 allowlist schema로 생성한다.
  • NOT_CAPTURED placeholder는 qualification 성공으로 인정하지 않는다.
  • sanitizer/reconciler 성공 뒤에도 human clean commit/remote CI 전에는 releaseQualification=NOT_CLAIMED를 유지한다.
  • 성공/실패 모두 exact VM allowlist를 teardown하고 lab resource가 0인지 확인한다. local retain-on-failure opt-in은 명시된 경우만 허용하고 CI에서는 금지한다.

Task 12 — CI, runbook, verification와 Wiki capture

CI

  • PR blocking redis-standalone job을 release-gate.needs와 result loop에 실제 포함한다.
  • nightly/release Redis production readiness workflow를 추가한다.
  • workflow contract test로 blocking job/aggregator 집합 동등성을 검증한다.

Verification

cd src
./gradlew :application-core:redisPolicyContractTest --console=plain
./gradlew :shared-contract:edgeRateLimitContractTest --console=plain
./gradlew :adapter:outbound:cache-redis:check --console=plain
./gradlew :app-bootstrap:redisCompositionTest --console=plain
./gradlew redisProductionReadiness --console=plain
./gradlew test --console=plain
./gradlew check --console=plain
./gradlew verifyCleanArchitectureDependencies --console=plain
./gradlew verifyPublicPathSnapshot --console=plain
./gradlew verifyEnvKeys --console=plain

Documentation

  • capability별 실제 readiness와 남은 R3 한계를 README/spec/runbook에 동기화한다.
  • 실행 명령, image/config/program digest, 실패/차단을 public LLM Wiki /home/donghyeon/workspace/ai-tools/llm-wiki/raw/branch-notes/main.md에 기록하고 실제 파생 오류/면접/블로그 raw 문서를 양방향 링크한다.

Completion gate

  • Task 11의 exit gate를 통과하면 Sentinel-first R2-ready candidate라고만 보고한다.
  • clean committed source와 실제 remote CI가 없으면 selected/R2로 승격하지 않는다.
  • 이 milestone 보고 뒤 멈추고 Cluster/R3/fenced coordination 또는 fileserver/HTTP client 중 다음 작업을 사용자와 다시 정한다.

Task 13 — Resume blocker: selection-driven role activation과 default boot

Problem

  • provider definition뿐 아니라 role binding도 capability가 선택되지 않으면 inert여야 한다.
  • 현재 구현은 role binding 전체를 runtime으로 열고 health contributor도 role property 존재만으로 활성화한다.
  • local 기본값에서 inbound rate-limit은 provider 없이 활성화되면 안 된다.

Tests first

  • CACHE/COORDINATION/SESSION deployment와 role을 모두 사전 선언해도 cache/rate/idempotency/lease/ session capability가 비활성이면 credential/trust resolution, native client, scheduler/subscriber, Redis health contributor가 모두 0이다.
  • 각 capability가 redis를 선택할 때만 해당 role이 활성화된다.
  • 같은 role을 쓰는 coordination capability 둘 이상은 하나의 runtime만 공유한다.
  • 선택 capability의 role binding이 빠지면 material resolution 전에 startup이 실패한다.
  • shipped .env와 실제 application.yml은 transport disabled/provider disabled 조합으로 기동 가능하고 중복 legacy rate-limit block이 없다.

Implementation

  • deployment/role registry validation과 runtime activation을 분리한다.
  • selectedCapabilities가 비어 있는 role은 registry/router/health에서 제외한다.
  • bootstrap health condition도 role property가 아니라 effective selected capability로 판단한다.
  • provider 설정은 inert 후보로 남기되 선택된 capability의 잘못된 role은 fail closed 한다.

Task 14 — Resume blocker: capability-aware semantic readiness

Problem

  • PING만으로 AVAILABLE/PROBE_SUCCEEDED를 선언하지 않는다.
  • required coordination/session은 실제 선택 capability의 program ACL과 최소 read/write 계약이 동작해야 ready다.

Tests first

  • PING은 성공하지만 SCRIPT LOAD/EVALSHA가 ACL로 거절된 coordination/session user는 redisRequired=DOWN이다.
  • capability별 representative program의 실제 key count와 command-to-key mapping을 그대로 검증한다. rate-limit의 state/dedup/order key와 session tombstone key 중 하나만 ACL pattern에서 빠져도 semantic readiness는 실패한다.
  • Redis 7.2 미만 server는 metadata 표기만으로 통과하지 않고 bounded runtime handshake에서 sanitized unsupported-version 상태가 된다.
  • 대표 program과 ACL probe script가 이미 warm인 상태에서도 runtime user의 SCRIPT LOAD 권한 누락을 별도로 탐지한다.
  • cache optional role에서 semantic probe 실패는 application liveness/readiness를 내리지 않고 DEGRADED만 보고한다.
  • 선언된 optional cache가 cold-start connect/PING에 일시 실패해도 context는 bounded unavailable route로 시작하고, health-triggered bounded single-flight reconnect 뒤 재시작 없이 복구한다. invalid configuration/material/program/schema는 계속 startup failure이며 required coordination/session은 fail closed다.
  • probe는 raw key/value, credential, server exception을 health detail에 노출하지 않는다.
  • probe key는 bounded, namespaced, TTL이 있고 성공/실패 후 잔여 상태가 없다.
  • saturation/recent command failure/closed route를 distinct sanitized reason으로 분류한다.
  • health scrape는 role별 minimum cadence와 single-flight로 full semantic suite 실행을 제한하고, cached observation의 시각/age를 노출해 stale success를 숨기지 않는다.

Implementation

  • role별 선택 capability를 입력으로 immutable semantic probe plan을 만든다.
  • probe는 catalog-owned bounded program과 capability-safe ephemeral operation만 사용한다.
  • optional cold-start outage는 resource-free unavailable runtime과 bounded on-demand reconnect로 표현하며 별도 unbounded scheduler/thread를 만들지 않는다. L1 invalidation subscription은 route recovery 시 실제 runtime에 다시 연결된다.
  • eviction은 runtime CONFIG 권한을 열지 않고 CONFIGURED_EXPECTATION_ONLY로 유지하며 외부 attestation 미완료를 readiness detail에 명시한다.

Task 15 — Resume blocker: bounded common primitive catalog

Problem

  • Deep design §14.6–§14.9의 자주 쓰는 race-safe helper가 아직 compare/delete 중심 R0 foundation에 머물러 있다.

Tests first

  • String, counter, hash, set, sorted-set, list baseline은 typed/versioned key, value/count/byte/deadline, role, slot, TTL, certainty bound를 강제한다.
  • bitmap/HLL/geo는 billing/auth correctness에 사용할 수 없는 explicit semantic classification과 offset/result/fan-in bound를 강제한다.
  • INCR -> EXPIRE, set/list admission, revision-CAS는 실제 Redis concurrency에서 atomic하다.
  • unbounded HGETALL, SMEMBERS, LRANGE, arbitrary command/script surface는 제공하지 않는다.

Implementation

  • package-private RedisPrimitiveCatalog과 structure별 bounded facade를 Redis leaf 내부에 둔다.
  • application/shared public API에는 Redis command나 raw key를 노출하지 않는다.
  • 아직 실제 semantic consumer가 없는 primitive는 Spring bean/public capability로 노출하지 않는다.

Task 16 — Resume blocker: capability observability와 graceful lifecycle

Tests first

  • cache/rate/idempotency/lease/session의 operation, outcome, certainty, role, queue/latency가 bounded low-cardinality metric/event로 관측된다.
  • raw key, subject, session/idempotency/lease token, secret reference/value, exception message는 tag/log/trace에 들어가지 않는다.
  • optional cache와 required coordination/session의 failure signal이 health와 metric에서 일치한다.
  • shutdown은 subscriber/scheduler/router/runtime 순서로 bounded drain되고 새 command를 거절한다.

Implementation

  • framework-neutral observation event/port와 Micrometer rendering을 계층 소유권에 맞게 둔다.
  • trace/log는 기존 skeleton observability 경계를 재사용하고 Redis native type을 core에 유출하지 않는다.
  • docs/registries/metrics.yaml과 runbook을 실제 emitted metric과 동기화한다.

Task 17 — Resume final review, readiness truth, verification와 Wiki

  • Task 1316을 task별 spec/code-quality review한다.
  • Redis deep design §39/§40을 독립 재검토해 selected/implemented-candidate/not-implemented를 실제 evidence와 일치시킨다.
  • Sentinel/Cluster/k3s/R3 evidence가 없으면 지원/완료로 표기하지 않는다.
  • Task 12의 전체 검증을 실행하고 동시 작업의 비-Redis 실패는 소유 파일과 증거를 분리한다.
  • Redis README/spec/runbook, readiness registry, CI artifact 계약을 동기화한다.
  • LLM Wiki branch-note와 실제 파생 raw 문서를 양방향 링크로 캡처한다.