663 lines
37 KiB
Markdown
663 lines
37 KiB
Markdown
# Redis Production Capability Completion Plan
|
||
|
||
> **Scope:** Redis를 먼저 완료한다. 현재 실행 단위는 deep design Phase 5 전체가 아니라
|
||
> `Sentinel-first R2 qualification slice`다. 이 slice의 검증과 보고가 끝나면 멈추고
|
||
> fileserver, HTTP client, Redis Cluster/R3 중 다음 우선순위를 다시 정한다.
|
||
>
|
||
> **Workflow note:** 저장소가 지정한 Superpowers 설계·계획·TDD·디버깅·검증·리뷰 워크플로우를
|
||
> 적용한다. agent는 human-only commit 정책에 따라 stage/commit/amend/push하지 않는다.
|
||
|
||
**Goal:** `2026-07-26-redis-production-capability-design.md`의 Phase 1–5를 capability별로 구현하고,
|
||
standalone 기능의 존재를 production readiness로 오표기하지 않는 Redis platform을 만든다.
|
||
|
||
**Architecture:** `application-core`와 `shared-contract`는 provider-neutral semantic contract만
|
||
소유한다. `adapter:outbound:cache-redis`가 Redis deployment, topology, key, codec, program,
|
||
runtime과 capability provider를 소유한다. `adapter:inbound:web`은 HTTP rate/session 보안 매핑만,
|
||
`app-bootstrap`은 provider/role/auth-mode composition만 소유한다. `domain-core`에는 Redis 개념을
|
||
추가하지 않는다.
|
||
|
||
**Readiness rule:** Redis leaf 전체에 단일 R2 label을 부여하지 않는다. `redis-cache`,
|
||
`redis-edge-rate-limit`, `redis-request-replay-idempotency`,
|
||
`redis-cache-refresh-soft-lease`, `redis-fenced-coordination`, `redis-session` card가 독립적으로
|
||
승격한다. R3 증거가 없는 failover/reshard/rotation은 R2 범위로 과장하지 않는다.
|
||
|
||
**Worktree rule:** 현재 `main` worktree의 다른 기술 변경은 사용자 소유다. Redis가 소유하지 않는
|
||
fileserver, HTTP client, messaging, notification, object storage 변경을 되돌리거나 포맷하지 않는다.
|
||
|
||
**Current milestone exit:** agent-side 목표는 `R2-ready candidate`다. clean committed source와
|
||
실제 remote GitHub Actions evidence가 없으면 card를 `selected`로 바꾸거나 R2라고 주장하지 않는다.
|
||
|
||
---
|
||
|
||
## Task 0 — Baseline과 acceptance registry 고정
|
||
|
||
**Files**
|
||
|
||
- Create: `src/config/redis/readiness-cards.yaml`
|
||
- Create: `src/gradle/redis-test-images.properties`
|
||
- Modify: `src/adapter/outbound/cache-redis/README.md`
|
||
- Modify: `docs/superpowers/specs/2026-07-26-redis-production-capability-design.md`
|
||
|
||
**Tests first**
|
||
|
||
- registry가 canonical card ID 여섯 개를 정확히 한 번 포함하는지 실패 테스트를 작성한다.
|
||
- image tag에 exact version과 digest가 없으면 configuration이 실패하는 테스트를 작성한다.
|
||
- `selected`, `implemented-candidate`, `not-implemented` 이외 상태를 거절한다.
|
||
- 현재 구현과 다른 readiness 표기를 거절한다.
|
||
|
||
**Implementation**
|
||
|
||
- 시작 상태는 cache/rate를 `implemented-candidate`, 나머지는 `not-implemented`로 기록한다.
|
||
- 실제 required evidence가 생기기 전에는 어떤 card도 `selected` R2로 승격하지 않는다.
|
||
- Redis minimum version은 실행 가능한 image/digest와 program manifest를 한 SSOT로 맞춘다.
|
||
|
||
**Verification**
|
||
|
||
```bash
|
||
cd src
|
||
./gradlew :adapter:outbound:cache-redis:test --tests '*RedisReadinessRegistryTest' --console=plain
|
||
```
|
||
|
||
## Task 1 — Canonical deployment/topology/role model
|
||
|
||
**Files**
|
||
|
||
- Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderProperties.java`
|
||
- Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettings.java`
|
||
- Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactory.java`
|
||
- Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRole.java`
|
||
- Create: `src/adapter/outbound/cache-redis/src/main/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisRoleBinding.java`
|
||
- Test: `src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisDeploymentSettingsFactoryTest.java`
|
||
- Test: `src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/config/RedisProviderPropertiesBindingTest.java`
|
||
|
||
**Tests first**
|
||
|
||
- topology는 `standalone|sentinel|cluster` 중 정확히 하나다.
|
||
- endpoint는 non-empty, unique, bounded host/port다.
|
||
- Sentinel은 master name, 최소 3개 discovery endpoint, data/Sentinel auth와 TLS를 분리한다.
|
||
- Cluster는 database 0만 허용하고 seed가 비어 있으면 실패한다.
|
||
- role은 존재하는 deployment만 참조한다.
|
||
- cache와 session/coordination의 incompatible co-location을 startup 전에 거절한다.
|
||
- provider 정의만 있고 capability binding이 없으면 runtime side effect가 0이다.
|
||
|
||
**Implementation**
|
||
|
||
- Spring binding class와 validated sealed runtime model을 분리한다.
|
||
- legacy `app.cache.redis`와 `app.rate-limit`은 migration compiler 입력으로만 허용하고 canonical
|
||
model과 동시에 설정되면 precedence를 정하지 않고 실패한다.
|
||
- `ClientMode.EXTERNAL`을 topology로 취급하지 않는다.
|
||
|
||
**Verification**
|
||
|
||
```bash
|
||
cd src
|
||
./gradlew :adapter:outbound:cache-redis:test --tests '*RedisDeploymentSettings*' --console=plain
|
||
```
|
||
|
||
## Task 2 — Topology-aware runtime, TLS/ACL과 secret material
|
||
|
||
**Files**
|
||
|
||
- Create: `.../redis/runtime/RedisDeploymentRuntime.java`
|
||
- Create: `.../redis/runtime/RedisDeploymentRuntimeFactory.java`
|
||
- Create: `.../redis/runtime/StandaloneRedisDeploymentRuntime.java`
|
||
- Create: `.../redis/runtime/SentinelRedisDeploymentRuntime.java`
|
||
- Create: `.../redis/runtime/ClusterRedisDeploymentRuntime.java`
|
||
- Create: `.../redis/security/RedisCredentialMaterialProvider.java`
|
||
- Create: `.../redis/security/RedisCredentialRotationCoordinator.java`
|
||
- Modify: `src/adapter/outbound/cache-redis/build.gradle`
|
||
- Modify: `src/adapter/outbound/cache-redis/gradle.lockfile`
|
||
|
||
**Tests first**
|
||
|
||
- standalone/Sentinel/Cluster가 각자 다른 native client/runtime을 만든다.
|
||
- Sentinel discovery credential/trust와 data-node credential/trust가 섞이지 않는다.
|
||
- Cluster client는 periodic+adaptive topology refresh, DB 0, bounded redirect/queue profile을 가진다.
|
||
- production profile에서 plaintext, trust-all, hostname verification off를 거절한다.
|
||
- named ACL username이 없거나 raw password가 YAML에 있으면 production activation이 실패한다.
|
||
- duplicate/out-of-order rotation event, expiry 재조회, new connection 검증 실패가 old traffic을
|
||
안전하게 보존한다.
|
||
- disabled capability는 client/event-loop/subscriber/scheduler를 만들지 않는다.
|
||
|
||
**Implementation**
|
||
|
||
- direct `spring-data-redis`, `lettuce-core` dependency를 leaf가 소유한다.
|
||
- deployment별 client resources와 lifecycle을 소유한다.
|
||
- connect/TLS/acquire/command/overall/shutdown timeout을 분리한다.
|
||
- 기존 no-replay, disconnected reject, finite queue/count/byte admission을 topology runtime에도
|
||
보존한다.
|
||
- secret value/reference/provider exception을 log/metric에 남기지 않는다.
|
||
|
||
## Task 3 — Key, codec, program manifest foundation
|
||
|
||
**Files**
|
||
|
||
- Create: `src/config/redis/program-set.schema.json`
|
||
- Modify: `src/adapter/outbound/cache-redis/src/main/resources/redis/program-set.json`
|
||
- Modify: `src/adapter/outbound/cache-redis/src/main/resources/redis/rate-program-set.json`
|
||
- Modify: `.../redis/RedisProgramDescriptor.java`
|
||
- Modify: `.../redis/RedisProgramCatalog.java`
|
||
- Modify: `.../redis/RedisLuaProgramExecutor.java`
|
||
- Create: `.../redis/key/RedisKeyMaterialProvider.java`
|
||
- Create: `.../redis/codec/RedisCapabilityCodec.java`
|
||
|
||
**Tests first**
|
||
|
||
- 모든 program은 exact source digest, semantic version, ordered KEYS/ARGV, result schema, slot rule,
|
||
state/TTL bound, minimum Redis version, retry/certainty, ACL command를 가진다.
|
||
- manifest와 Java descriptor가 drift하면 build가 실패한다.
|
||
- `NOSCRIPT` recovery는 bounded `SCRIPT LOAD -> EVALSHA`이고 arbitrary source 실행 surface가 없다.
|
||
- same-resource multi-key는 real `CLUSTER KEYSLOT`과 같은 slot이다.
|
||
- key digest material rotation은 fixed/dual-read-delete/cold-cutover rule을 지킨다.
|
||
- cache/idempotency/session codec은 N/N-1, future/corrupt/oversize/forbidden type을 구분한다.
|
||
|
||
**Implementation**
|
||
|
||
- foundation/rate manifest를 하나의 versioned registry contract로 통합하되 capability package와
|
||
facade는 분리한다.
|
||
- raw command, raw key, generic program executor를 Spring/application public surface에 노출하지 않는다.
|
||
|
||
## Task 4 — Cache consistency spine와 semantic region composition
|
||
|
||
**Files**
|
||
|
||
- Modify: `src/application-core/src/main/java/dev/caskeleton/application/cache/*`
|
||
- Create: `.../redis/cache/RedisCacheGenerationStore.java`
|
||
- Create: `.../redis/cache/RedisCacheRegionCompiler.java`
|
||
- Add resources: `region-generation-init-v1.lua`, `region-generation-bump-v1.lua`,
|
||
`cache-record-if-generation-v1.lua`
|
||
- Modify: `.../redis/RedisStringCacheRegion.java`
|
||
- Tests: application barrier tests, Redis real-service concurrency tests, binding tests
|
||
|
||
**Tests first**
|
||
|
||
- source load 중 generation bump가 일어나면 old result가 visible하지 않다.
|
||
- captured generation과 source revision이 바뀌면 stale writer가 새 값을 덮어쓰지 않는다.
|
||
- generation init race에서 하나의 canonical generation만 선택된다.
|
||
- operation ID가 같은 bump replay는 한 번만 적용된다.
|
||
- 여러 semantic region의 duplicate/missing binding은 fail-fast다.
|
||
- 실제 consumer가 semantic `CacheRegionPort`와 `CacheAsideExecutor`를 사용하고 legacy fail-open
|
||
router와 암묵적으로 섞이지 않는다.
|
||
|
||
**Implementation decision**
|
||
|
||
- source revision은 opaque하므로 lexical “newer” 비교를 하지 않는다.
|
||
- region generation은 mass invalidation fence다.
|
||
- per-key invalidation은 해당 key의 revision/tombstone fence를 사용해 region 전체를 bump하지 않는다.
|
||
- write는 captured generation/revision condition을 만족할 때만 기록한다.
|
||
|
||
## Task 5 — Distributed refresh soft lease, L1/L2와 cache observability
|
||
|
||
**Files**
|
||
|
||
- Create application cache refresh coordination contracts without Redis types.
|
||
- Create Redis refresh claim/release programs and semantic provider.
|
||
- Create bounded L1 cache decorator and invalidation subscriber/reconciler.
|
||
- Create framework-free cache observation events and Micrometer adapter instrumentation.
|
||
- Update `docs/registries/metrics.yaml`.
|
||
|
||
**Tests first**
|
||
|
||
- 두 pod simulation에서 정상 시 refresh owner는 하나다.
|
||
- lease expiry에서는 duplicate load를 허용하지만 generation guard가 stale write를 차단한다.
|
||
- disconnected invalidation subscriber는 L1을 flush하고 generation을 재확인한다.
|
||
- Pub/Sub event loss에도 L1 TTL/generation reconciliation으로 stale bound를 지킨다.
|
||
- L1 max weight/cardinality/TTL, subscriber queue, refresh scheduler가 모두 bounded다.
|
||
- Redis liveness는 애플리케이션 liveness를 내리지 않는다.
|
||
- optional cache outage는 `DEGRADED`, required coordination/session outage는 `NOT_READY`다.
|
||
- cache role eviction/OOM에서 source concurrency와 queue가 bounded다.
|
||
|
||
## Task 6 — Edge rate limit end-to-end
|
||
|
||
**Files**
|
||
|
||
- Modify: `src/shared-contract/src/main/java/dev/caskeleton/shared/ratelimit/*`
|
||
- Modify: `src/adapter/inbound/web/src/main/java/dev/caskeleton/adapter/inbound/web/ratelimit/*`
|
||
- Modify: `src/adapter/outbound/cache-redis/src/main/java/.../redis/*rate*`
|
||
- Modify: `src/app-bootstrap` composition
|
||
|
||
**Tests first**
|
||
|
||
- inbound가 process-local map이 아니라 `EdgeRateLimitPort`를 호출한다.
|
||
- subject는 raw principal/IP가 아닌 bounded pseudonymous digest다.
|
||
- fixed/sliding-counter/token-bucket reference/property/concurrency vector를 통과한다.
|
||
- evaluation ID replay가 quota를 두 번 소비하지 않는다.
|
||
- bounded local emergency는 configured degraded provider일 때만 동작한다.
|
||
- Redis/local/disabled provider exclusivity, shadow/degraded source, 429/503와 `Retry-After` mapping을
|
||
검증한다.
|
||
- legacy unbounded map과 silent primary fallback을 제거한다.
|
||
|
||
## Task 7 — Idempotency v2와 Redis provider
|
||
|
||
**Files**
|
||
|
||
- Replace/extend `src/application-core/.../idempotency` with owner-safe v2 contracts.
|
||
- Add Redis idempotency state programs/provider/codec.
|
||
- Migrate the existing JPA provider to the same semantic contract only after checking its separate
|
||
worktree changes; never overwrite concurrent persistence work.
|
||
|
||
**Tests first**
|
||
|
||
- atomic claim, fingerprint mismatch, owner/attempt-safe start/renew/complete/fail/release/inspect.
|
||
- processing TTL과 replay TTL 분리.
|
||
- expired `CLAIMED` takeover, expired `EXECUTING -> RECOVERY_REQUIRED`.
|
||
- response-loss replay/reconciliation, conflicting response digest reject.
|
||
- unverified cross-store effect는 자동 discard/re-execution하지 않는다.
|
||
- JDBC/Redis provider가 같은 scope를 동시에 claim하지 않는다.
|
||
|
||
**Implementation**
|
||
|
||
- Redis가 cross-store exactly-once를 보장한다고 표현하지 않는다.
|
||
- JPA migration 충돌이 있으면 Redis completion의 명시적 integration blocker로 보고하고 해당
|
||
worktree의 결과와 재대조한다.
|
||
|
||
## Task 8 — Efficiency lease와 optional fenced coordination
|
||
|
||
**Tests first**
|
||
|
||
- acquire/inspect/renew/release가 owner+operation token을 비교한다.
|
||
- response loss는 `UNKNOWN/INDETERMINATE`이며 same token inspect로 reconcile한다.
|
||
- expired old owner는 renew/release할 수 없다.
|
||
- watchdog는 bounded scheduler와 cancellation을 사용하고 lost 상태를 전달한다.
|
||
- fenced card를 선택하면 durable epoch/high-watermark 등록과 protected-resource stale-token reject를
|
||
실제 fixture로 증명한다.
|
||
|
||
**Implementation**
|
||
|
||
- close-only `DistributedLock`은 compatibility facade로 유지하되 새 코드가 strong lock으로
|
||
오해하지 않게 guarantee를 명명한다.
|
||
- fencing 없는 Redis lease를 business correctness lock으로 광고하지 않는다.
|
||
|
||
## Task 9 — Redis Session과 JWT/session exclusive composition
|
||
|
||
**Files**
|
||
|
||
- Add direct `spring-session-core` and `spring-session-data-redis` to Redis leaf.
|
||
- Add adapter-internal versioned session store/programs/serializer.
|
||
- Add inbound web cookie/CSRF/fixation settings and security configuration.
|
||
- Add app-bootstrap `jwt|redis-session` exclusive composition.
|
||
|
||
**Tests first**
|
||
|
||
- JWT mode는 session Redis connection/bean/thread side effect가 0이다.
|
||
- pod A create/save, pod B read/touch/logout.
|
||
- idle/absolute expiry, rotation, old ID reject, stale save after logout reject.
|
||
- explicit allowlisted serializer N/N-1 and corrupt payload re-auth.
|
||
- secure/httpOnly/SameSite/host-only cookie, CSRF enabled, fixation rotation.
|
||
- repository outage/noeviction OOM/failover는 fail-open 인증으로 바뀌지 않는다.
|
||
- indexed repository는 별도 opt-in이며 Cluster event cleanup 한계를 독립 검증한다.
|
||
|
||
## Task 10 — Real-service, topology, fault와 readiness Gradle tasks
|
||
|
||
**Files**
|
||
|
||
- Create: `src/adapter/outbound/cache-redis/src/redisTest/**`
|
||
- Modify: `src/adapter/outbound/cache-redis/build.gradle`
|
||
- Modify: `src/build.gradle`
|
||
- Create/update Redis test topology resources and sanitized evidence reporter
|
||
|
||
**Public tasks**
|
||
|
||
- `redisStandaloneTest`, `redisSecurityTest`, `redisSentinelTest`, `redisClusterTest`,
|
||
`redisFaultTest`, `redisCompatibilityTest`
|
||
- capability card test/readiness tasks named exactly as Redis deep design §37.22
|
||
- root `redisProductionReadiness`, `redisAllImplementedCandidates`
|
||
|
||
**Rules**
|
||
|
||
- selected evidence에서 Docker/service 부재나 0 discovered tests는 failure다.
|
||
- unselected card는 skipped가 아니라 `not selected`다.
|
||
- image/program/config digest와 sanitized JUnit/topology timeline을 evidence artifact로 남긴다.
|
||
|
||
## Task 11 — Container topology와 3-node k3s qualification
|
||
|
||
이번 실행은 deep design §37.13/Phase 5A의 Sentinel-first slice만 다룬다. Cluster, fenced
|
||
coordination, R3 long chaos/soak, k3s control-plane HA, physical host/AZ failure, full
|
||
credential/certificate rotation은 후속 작업이다.
|
||
|
||
### Task 11.1 — Lab lifecycle contract와 host isolation RED
|
||
|
||
이 작업은 리뷰 경계를 다음처럼 분리한다. 두 하위 작업이 모두 독립 리뷰를 통과하기 전에는 부모
|
||
Task 11.1을 완료로 표시하지 않는다.
|
||
|
||
- `Task 11.1A-1`: VM lifecycle, ownership marker/state, lock/signal/handoff cleanup, host
|
||
fingerprint와 bounded command. 현재 구현을 동결한다.
|
||
- `Task 11.1A-2`: pinned K3s generated-kubeconfig strict validator/renderer. 실행 계획은
|
||
`docs/superpowers/plans/2026-07-30-redis-lab-strict-kubeconfig-renderer.md`를 따른다.
|
||
|
||
2026-07-30 상태: `Task 11.1A-1` lifecycle/ownership과 `Task 11.1A-2` strict renderer는
|
||
whole-task 독립 review에서 Critical `0`, Important `0`, Minor `0`, SPEC PASS /
|
||
QUALITY APPROVED를 받았다. fresh direct/Gradle fake-only 검증도 통과해 부모 `Task 11.1A`의
|
||
fake-only 범위는 완료다. 이는 live VM/k3s/kubectl/network/host qualification이나 Redis
|
||
R2 readiness 완료를 의미하지 않는다.
|
||
|
||
**Tracked files**
|
||
|
||
- Create: `infra/redis-lab/README.md`
|
||
- Create: `infra/redis-lab/versions.env`
|
||
- Create: `infra/redis-lab/bin/redis-lab`
|
||
- Create: `infra/redis-lab/cloud-init/node.yaml`
|
||
- Create: `infra/redis-lab/test/redis-lab-contract.sh`
|
||
- Modify: Redis Gradle VM-free lifecycle contract task
|
||
|
||
**Tests first**
|
||
|
||
- VM 이름은 `ca-redis-lab-server`, `ca-redis-lab-agent-1`,
|
||
`ca-redis-lab-agent-2` exact allowlist만 허용한다.
|
||
- server 1 + agent 2, resource `2/3GiB/12GiB`, `2/2.5GiB/12GiB`,
|
||
`2/2.5GiB/12GiB`, pod CIDR `10.52.0.0/16`, service CIDR
|
||
`10.53.0.0/16`, context `ca-redis-lab`을 검증한다.
|
||
- host 관측은 default kubeconfig의 run-scoped copy와 원래 host context를 사용하고 read-only
|
||
allowlist만 허용한다. lab 호출은 별도 ignored `src/build/redis-lab/kubeconfig`와 exact
|
||
`ca-redis-lab` context를 사용한다.
|
||
- default kubeconfig merge/write, host context mutation, wildcard VM cleanup, global
|
||
`multipass purge`를 정적/동적 contract가 거절한다.
|
||
- preflight/postflight host kubeconfig/context/node/workload fingerprint가 다르면 실패한다.
|
||
- CI는 retain-on-failure를 거절하고, local opt-in만 exact VM 보존을 허용한다.
|
||
- fake `multipass`/`kubectl`을 주입하는 shell contract는 partial-create cleanup과 exact command
|
||
allowlist를 VM 생성 없이 검증하고 `redisLabContractTest`로 module `check`에 연결한다.
|
||
- launch 전 exact name을 run-owned `PENDING`으로 atomic 예약하고 성공 직후 `CREATED`로
|
||
승격한다. timeout/실패/상태 승격 실패는 이 run이 예약한 exact name만 정리한다.
|
||
- private run-scoped rendered cloud-init은 non-secret `RUN_ID|VM_NAME` ownership marker를
|
||
기록한다. cleanup/down은 bounded marker read가 state owner와 exact name 일치를 증명할
|
||
때만 delete한다. launch timeout/error는 `RECONCILE` tombstone과 bounded late-create poll로
|
||
처리하며 absent/unreadable/mismatch는 delete/state removal 없이 fail-closed한다.
|
||
- lifecycle 전체는 nonblocking exclusive lock과 run identity를 사용한다. direct `up`과
|
||
`run` 모두 첫 launch 전 emergency cleanup을 활성화하고, signal/concurrent 실행이 다른
|
||
run state나 VM을 채택·삭제하지 못한다. user command에는 lock file descriptor를 상속하지
|
||
않으며 기본 bounded external child도 FD를 닫고 lock acquisition만 예외로 유지한다.
|
||
`run`의 inner `up` 성공과 user command 시작 사이에도 cleanup-required flag가 연속 유지돼
|
||
zero-ownership handoff gap이 없어야 한다.
|
||
- host kubeconfig copy는 fingerprint/CIDR 관측 범위가 끝나면 성공/실패와 무관하게 제거한다.
|
||
- lab kubeconfig renderer는 denylist/generic-count 보강을 사용하지 않는다. pinned K3s의
|
||
canonical block-style one-cluster/context/user grammar를 별도 tracked AWK state machine으로
|
||
allowlist하며, catch-all pass-through 없이 duplicate/extra/reordered/unknown/flow-style
|
||
identity와 모든 비허용 구조를 fail-closed로 거절한다.
|
||
- external command와 3-node Ready 대기는 bounded이고, host service CIDR은 assigned
|
||
ClusterIP에서 추측하지 않고 명시적 validated input 또는 신뢰 가능한 host 설정에서 얻는다.
|
||
- mutable `curl | sudo sh` installer는 금지한다. exact K3s release URL과 SHA-256을 repository에
|
||
pin하고 host download와 각 VM transfer 뒤 다시 검증한 후에만 install/start한다.
|
||
- shell contract는 별도 fixture repository에서 실행하고 actual `src/build/redis-lab` canary를
|
||
byte-for-byte 보존한다. fake PATH는 explicit safe wrapper 외 모든 명령을 fail-closed한다.
|
||
|
||
### Task 11.2A — Sentinel manifest와 security static contract GREEN
|
||
|
||
**Tracked files**
|
||
|
||
- Create: `infra/redis-lab/config/redis.conf.tmpl`
|
||
- Create: `infra/redis-lab/config/sentinel.conf.tmpl`
|
||
- Create: `infra/redis-lab/config/redis-users.acl.tmpl`
|
||
- Create: `infra/redis-lab/config/sentinel-users.acl.tmpl`
|
||
- Create: `infra/redis-lab/k3s/namespace.yaml`
|
||
- Create: `infra/redis-lab/k3s/redis-data.yaml`
|
||
- Create: `infra/redis-lab/k3s/redis-sentinel.yaml`
|
||
- Create: `infra/redis-lab/k3s/network-policy.yaml`
|
||
- Create:
|
||
`src/adapter/outbound/cache-redis/src/test/java/dev/caskeleton/adapter/outbound/cache/redis/RedisLabManifestContractTest.java`
|
||
- Modify: Redis Gradle manifest contract task
|
||
|
||
- static contract와 live security evidence를 분리한다. YAML/템플릿 정적 통과는 TLS handshake,
|
||
ACL authorization, CNI enforcement, scheduling/failover의 실행 증거가 아니다.
|
||
- data Redis 3개와 Sentinel 3개는 각각 stable ordinal/headless DNS가 필요한 StatefulSet으로
|
||
구성하고 `kubernetes.io/hostname` required anti-affinity와 `maxSkew=1/DoNotSchedule`
|
||
topology spread, `podManagementPolicy: Parallel`을 적용한다.
|
||
- data는 PVC + AOF `appendfsync everysec`를 사용한다. Sentinel은 공식 동작상 writable config에
|
||
discovery/failover 상태를 rewrite하므로, bootstrap source를 pod별 writable PVC config로
|
||
최초 1회 atomic init-copy하고 restart 때 기존 rewritten config를 덮어쓰지 않는다.
|
||
비어 있거나 손상된 기존 config는 자동 복구로 덮지 않고 startup을 실패시킨다.
|
||
- Redis image SSOT는 `src/gradle/redis-test-images.properties`의
|
||
`redis.minimum.image` exact tag+digest다. `redis.approved.image`나 임의 YAML image를 이
|
||
minimum-version Sentinel slice에 섞지 않는다.
|
||
- plaintext port는 data/Sentinel 모두 0이고 TLS port만 연다. `tls-replication yes`,
|
||
hostname resolution/announcement와 certificate SAN용 stable DNS를 사용한다. data plane과
|
||
Sentinel plane은 서로 다른 CA/leaf material을 가지며, peer 연결에 필요한 root만 명시적
|
||
trust bundle로 교차 포함한다.
|
||
- ACL identity를 하나의 `redis-user`로 합치지 않는다.
|
||
- application data user: 선택 capability/program command/key/channel만;
|
||
- replica user: `+psync +replconf +ping`;
|
||
- Sentinel-to-data user: 공식 최소 Sentinel control command/channel set;
|
||
- Sentinel peer user: Sentinel 간 통신에 필요한 동일 superuser credential;
|
||
- application Sentinel discovery user: auth/hello/ping/role과 allowlisted read-only
|
||
`SENTINEL` subcommand만.
|
||
default user는 off이며 application/data/discovery user에 `+@all`, `allkeys`,
|
||
`allchannels`를 주지 않는다.
|
||
- Redis data ACL과 Sentinel ACL은 별도 template/projection이다. Sentinel peer superuser가
|
||
data Redis에, data capability user가 Sentinel에 존재하면 static contract가 실패한다.
|
||
- Secret/CA/private key/rendered config는 run별 `umask 077` 아래 생성하고 tracked manifest에는
|
||
Secret value, PEM, password가 없다. probe/command line에 `--pass`를 쓰지 않는다.
|
||
- exec probe를 사용해 kubelet source CIDR 예외를 만들지 않는다. default-deny ingress/egress
|
||
뒤 data 6379, Sentinel 26379, kube-dns와 exact qualification/application pod selector만
|
||
허용한다.
|
||
- Service는 headless/ClusterIP만, PDB는 data/Sentinel 각각 `minAvailable: 2`, container는
|
||
non-root, read-only root filesystem, privilege escalation false, capabilities drop ALL,
|
||
seccomp RuntimeDefault, explicit requests/limits를 요구한다.
|
||
- structural positive test와 한 필드씩 제거/변조한 mutation-negative fixture가
|
||
anti-affinity, spread, PDB, probes, TLS-only, ACL separation, Secret reference,
|
||
NetworkPolicy, image SSOT를 실제로 fail시키는지 검증한다.
|
||
- `hostPath`, `hostNetwork`, `hostPID`, `hostIPC`, privileged, NodePort, LoadBalancer,
|
||
tracked Secret data/stringData/PEM과 implicit latest image를 거절한다.
|
||
- static validator는 exact document inventory, duplicate YAML key/identity, selector/template
|
||
일치, exact NetworkPolicy edge graph를 검증한다. 정적 ordinal bootstrap은 최초
|
||
`redis-data-0` primary와 두 replica만 증명하며, failover 뒤 old-primary 재합류와 stale
|
||
direct write 차단은 live gate에 남긴다.
|
||
|
||
### Task 11.2B — Sentinel workload와 live security baseline GREEN
|
||
|
||
- Redis primary 1 + replica 2와 Sentinel 3/quorum 2를 세 node에 분산한다.
|
||
- anti-affinity/topology spread, PDB, NetworkPolicy, separate data/Sentinel CA와 named ACL을
|
||
적용한다.
|
||
- secret/certificate/k3s token은 매 run `umask 077` transient material로 생성하고 tracked
|
||
manifest에는 값/PEM을 넣지 않는다. Sentinel bootstrap config는 Secret volume에서 pod별
|
||
writable PVC로 최초 1회 atomic init-copy하며, 기존 rewritten config를 덮어쓰지 않는다.
|
||
- Redis image는 `redis.minimum.image` exact image/digest를 render하고 실제 pod image
|
||
ID/digest가 일치하는지 수집한다.
|
||
- data credential/CA로 Sentinel discovery가 실패하고 Sentinel material로 data command가
|
||
실패하는 negative test, untrusted CA/hostname mismatch/plaintext rejection을 실행한다.
|
||
- `SENTINEL CKQUORUM`, writable config rewrite/restart, exact 3 Ready placement, PDB,
|
||
default-deny/explicit-allow NetworkPolicy enforcement를 live k3s에서 검증한다.
|
||
- failover 중 죽어 있던 old primary가 재합류할 때 readiness가 stale direct write를 허용하지
|
||
않고 새 primary의 replica로 수렴하는지 live 검증한다.
|
||
|
||
### Task 11.3 — Sentinel client runtime TDD
|
||
|
||
- current `UnsupportedOperationException`을 먼저 고정하는 test를 quorum-consistent discovery와
|
||
분리된 discovery/data material contract로 교체한다.
|
||
- 2-of-3 Sentinel이 같은 primary를 보고할 때만 후보를 만들고 loopback/wildcard/unexpected
|
||
endpoint를 거절한다.
|
||
- active Sentinel role이 있을 때만 registry당 daemon worker 1개, role당 fixed-delay task 1개를
|
||
만들고 `sentinel-discovery-refresh-period`(기본 30초, 5초..5분)를 적용한다.
|
||
- scheduled poll과 command failure-triggered immediate rediscovery는 role별 같은 single-flight를
|
||
공유한다. `snapshot()`은 보조 trigger일 뿐 정상 polling을 대신하지 않는다.
|
||
- 정상 poll은 Sentinel material만 해석하고 현재 route identity와 같으면 data material/client를
|
||
만들지 않는다. 바뀐 quorum-approved endpoint에만 data candidate를 연다.
|
||
- command failure listener는 route lease 반환 뒤 topology/connectivity `UNAVAILABLE`에만
|
||
동작하며 listener 실패가 원래 certainty를 덮어쓰지 않는다.
|
||
- 새 data runtime은 version/program/semantic readiness를 통과한 뒤 router에 install한다.
|
||
- opaque route identity와 monotonic generation token으로 stale/same-primary candidate를
|
||
거절하고, install된 경우 old runtime은 new admission을 닫고 bounded drain/close한다.
|
||
- close는 task/worker를 bounded 종료하고 late candidate를 install하지 않고 정확히 한 번 닫는다.
|
||
- mutation을 자동 replay하지 않고 실행 여부가 불명확하면 `INDETERMINATE`를 보존한다.
|
||
|
||
### Task 11.4 — Multi-pod normal/failover qualification
|
||
|
||
1. host/lab preflight와 3 node/Sentinel quorum readiness를 수집한다.
|
||
2. 서로 다른 application pod에서 rate limit evaluation replay, idempotency
|
||
claim/start/renew/complete, session create/read/touch/rotate/revoke를 검증한다.
|
||
3. current primary pod를 kill하고 readiness unavailable timestamp를 기록한다.
|
||
4. Sentinel quorum election, client rediscovery, runtime generation swap/drain, semantic
|
||
readiness recovery를 실제 순서대로 기록한다.
|
||
5. election 60초, 추가 rediscovery/swap 30초, 총 recovery 90초의 regression limit을 적용한다.
|
||
6. rate state가 조용히 reset되지 않고 idempotency owner/terminal 결과가 중복되지 않으며
|
||
confirmed session state가 유지되는지 확인한다.
|
||
7. old primary의 replica 재합류와 모든 actor의 동일 generation 관측을 확인한다.
|
||
|
||
correctness role에는 bounded `min-replicas-to-write`/`min-replicas-max-lag`와 명시적 replica
|
||
acknowledgement policy를 사용한다. zero-data-loss/strong consistency를 주장하지 않으며
|
||
response-only cut 등 실행 여부가 불확실한 mutation은 `INDETERMINATE`이고 blind retry하지 않는다.
|
||
|
||
### Task 11.5 — Evidence와 exact teardown
|
||
|
||
- actual image digest/image ID, config/program digest, sanitized fault/election/recovery timeline,
|
||
capability별 outcome/certainty, Kubernetes/Sentinel 관측을 allowlist schema로 생성한다.
|
||
- `NOT_CAPTURED` placeholder는 qualification 성공으로 인정하지 않는다.
|
||
- sanitizer/reconciler 성공 뒤에도 human clean commit/remote CI 전에는
|
||
`releaseQualification=NOT_CLAIMED`를 유지한다.
|
||
- 성공/실패 모두 exact VM allowlist를 teardown하고 lab resource가 0인지 확인한다. local
|
||
retain-on-failure opt-in은 명시된 경우만 허용하고 CI에서는 금지한다.
|
||
|
||
## Task 12 — CI, runbook, verification와 Wiki capture
|
||
|
||
**CI**
|
||
|
||
- PR blocking `redis-standalone` job을 `release-gate.needs`와 result loop에 실제 포함한다.
|
||
- nightly/release Redis production readiness workflow를 추가한다.
|
||
- workflow contract test로 blocking job/aggregator 집합 동등성을 검증한다.
|
||
|
||
**Verification**
|
||
|
||
```bash
|
||
cd src
|
||
./gradlew :application-core:redisPolicyContractTest --console=plain
|
||
./gradlew :shared-contract:edgeRateLimitContractTest --console=plain
|
||
./gradlew :adapter:outbound:cache-redis:check --console=plain
|
||
./gradlew :app-bootstrap:redisCompositionTest --console=plain
|
||
./gradlew redisProductionReadiness --console=plain
|
||
./gradlew test --console=plain
|
||
./gradlew check --console=plain
|
||
./gradlew verifyCleanArchitectureDependencies --console=plain
|
||
./gradlew verifyPublicPathSnapshot --console=plain
|
||
./gradlew verifyEnvKeys --console=plain
|
||
```
|
||
|
||
**Documentation**
|
||
|
||
- capability별 실제 readiness와 남은 R3 한계를 README/spec/runbook에 동기화한다.
|
||
- 실행 명령, image/config/program digest, 실패/차단을 public LLM Wiki
|
||
`/home/donghyeon/workspace/ai-tools/llm-wiki/raw/branch-notes/main.md`에
|
||
기록하고 실제 파생 오류/면접/블로그 raw 문서를 양방향 링크한다.
|
||
|
||
**Completion gate**
|
||
|
||
- Task 11의 exit gate를 통과하면 `Sentinel-first R2-ready candidate`라고만 보고한다.
|
||
- clean committed source와 실제 remote CI가 없으면 selected/R2로 승격하지 않는다.
|
||
- 이 milestone 보고 뒤 멈추고 Cluster/R3/fenced coordination 또는 fileserver/HTTP client 중
|
||
다음 작업을 사용자와 다시 정한다.
|
||
|
||
## Task 13 — Resume blocker: selection-driven role activation과 default boot
|
||
|
||
**Problem**
|
||
|
||
- provider definition뿐 아니라 role binding도 capability가 선택되지 않으면 inert여야 한다.
|
||
- 현재 구현은 role binding 전체를 runtime으로 열고 health contributor도 role property 존재만으로
|
||
활성화한다.
|
||
- local 기본값에서 inbound rate-limit은 provider 없이 활성화되면 안 된다.
|
||
|
||
**Tests first**
|
||
|
||
- CACHE/COORDINATION/SESSION deployment와 role을 모두 사전 선언해도 cache/rate/idempotency/lease/
|
||
session capability가 비활성이면 credential/trust resolution, native client, scheduler/subscriber,
|
||
Redis health contributor가 모두 0이다.
|
||
- 각 capability가 `redis`를 선택할 때만 해당 role이 활성화된다.
|
||
- 같은 role을 쓰는 coordination capability 둘 이상은 하나의 runtime만 공유한다.
|
||
- 선택 capability의 role binding이 빠지면 material resolution 전에 startup이 실패한다.
|
||
- shipped `.env`와 실제 `application.yml`은 transport disabled/provider disabled 조합으로 기동
|
||
가능하고 중복 legacy rate-limit block이 없다.
|
||
|
||
**Implementation**
|
||
|
||
- deployment/role registry validation과 runtime activation을 분리한다.
|
||
- `selectedCapabilities`가 비어 있는 role은 registry/router/health에서 제외한다.
|
||
- bootstrap health condition도 role property가 아니라 effective selected capability로 판단한다.
|
||
- provider 설정은 inert 후보로 남기되 선택된 capability의 잘못된 role은 fail closed 한다.
|
||
|
||
## Task 14 — Resume blocker: capability-aware semantic readiness
|
||
|
||
**Problem**
|
||
|
||
- PING만으로 `AVAILABLE/PROBE_SUCCEEDED`를 선언하지 않는다.
|
||
- required coordination/session은 실제 선택 capability의 program ACL과 최소 read/write 계약이
|
||
동작해야 ready다.
|
||
|
||
**Tests first**
|
||
|
||
- PING은 성공하지만 `SCRIPT LOAD`/`EVALSHA`가 ACL로 거절된 coordination/session user는
|
||
`redisRequired=DOWN`이다.
|
||
- capability별 representative program의 실제 key count와 command-to-key mapping을 그대로
|
||
검증한다. rate-limit의 state/dedup/order key와 session tombstone key 중 하나만 ACL pattern에서
|
||
빠져도 semantic readiness는 실패한다.
|
||
- Redis 7.2 미만 server는 metadata 표기만으로 통과하지 않고 bounded runtime handshake에서
|
||
sanitized unsupported-version 상태가 된다.
|
||
- 대표 program과 ACL probe script가 이미 warm인 상태에서도 runtime user의 `SCRIPT LOAD`
|
||
권한 누락을 별도로 탐지한다.
|
||
- cache optional role에서 semantic probe 실패는 application liveness/readiness를 내리지 않고
|
||
`DEGRADED`만 보고한다.
|
||
- 선언된 optional cache가 cold-start connect/PING에 일시 실패해도 context는 bounded unavailable
|
||
route로 시작하고, health-triggered bounded single-flight reconnect 뒤 재시작 없이 복구한다.
|
||
invalid configuration/material/program/schema는 계속 startup failure이며 required
|
||
coordination/session은 fail closed다.
|
||
- probe는 raw key/value, credential, server exception을 health detail에 노출하지 않는다.
|
||
- probe key는 bounded, namespaced, TTL이 있고 성공/실패 후 잔여 상태가 없다.
|
||
- saturation/recent command failure/closed route를 distinct sanitized reason으로 분류한다.
|
||
- health scrape는 role별 minimum cadence와 single-flight로 full semantic suite 실행을 제한하고,
|
||
cached observation의 시각/age를 노출해 stale success를 숨기지 않는다.
|
||
|
||
**Implementation**
|
||
|
||
- role별 선택 capability를 입력으로 immutable semantic probe plan을 만든다.
|
||
- probe는 catalog-owned bounded program과 capability-safe ephemeral operation만 사용한다.
|
||
- optional cold-start outage는 resource-free unavailable runtime과 bounded on-demand reconnect로
|
||
표현하며 별도 unbounded scheduler/thread를 만들지 않는다. L1 invalidation subscription은
|
||
route recovery 시 실제 runtime에 다시 연결된다.
|
||
- eviction은 runtime `CONFIG` 권한을 열지 않고 `CONFIGURED_EXPECTATION_ONLY`로 유지하며 외부
|
||
attestation 미완료를 readiness detail에 명시한다.
|
||
|
||
## Task 15 — Resume blocker: bounded common primitive catalog
|
||
|
||
**Problem**
|
||
|
||
- Deep design §14.6–§14.9의 자주 쓰는 race-safe helper가 아직 compare/delete 중심 R0 foundation에
|
||
머물러 있다.
|
||
|
||
**Tests first**
|
||
|
||
- String, counter, hash, set, sorted-set, list baseline은 typed/versioned key, value/count/byte/deadline,
|
||
role, slot, TTL, certainty bound를 강제한다.
|
||
- bitmap/HLL/geo는 billing/auth correctness에 사용할 수 없는 explicit semantic classification과
|
||
offset/result/fan-in bound를 강제한다.
|
||
- `INCR -> EXPIRE`, set/list admission, revision-CAS는 실제 Redis concurrency에서 atomic하다.
|
||
- unbounded `HGETALL`, `SMEMBERS`, `LRANGE`, arbitrary command/script surface는 제공하지 않는다.
|
||
|
||
**Implementation**
|
||
|
||
- package-private `RedisPrimitiveCatalog`과 structure별 bounded facade를 Redis leaf 내부에 둔다.
|
||
- application/shared public API에는 Redis command나 raw key를 노출하지 않는다.
|
||
- 아직 실제 semantic consumer가 없는 primitive는 Spring bean/public capability로 노출하지 않는다.
|
||
|
||
## Task 16 — Resume blocker: capability observability와 graceful lifecycle
|
||
|
||
**Tests first**
|
||
|
||
- cache/rate/idempotency/lease/session의 operation, outcome, certainty, role, queue/latency가 bounded
|
||
low-cardinality metric/event로 관측된다.
|
||
- raw key, subject, session/idempotency/lease token, secret reference/value, exception message는
|
||
tag/log/trace에 들어가지 않는다.
|
||
- optional cache와 required coordination/session의 failure signal이 health와 metric에서 일치한다.
|
||
- shutdown은 subscriber/scheduler/router/runtime 순서로 bounded drain되고 새 command를 거절한다.
|
||
|
||
**Implementation**
|
||
|
||
- framework-neutral observation event/port와 Micrometer rendering을 계층 소유권에 맞게 둔다.
|
||
- trace/log는 기존 skeleton observability 경계를 재사용하고 Redis native type을 core에 유출하지
|
||
않는다.
|
||
- `docs/registries/metrics.yaml`과 runbook을 실제 emitted metric과 동기화한다.
|
||
|
||
## Task 17 — Resume final review, readiness truth, verification와 Wiki
|
||
|
||
- Task 13–16을 task별 spec/code-quality review한다.
|
||
- Redis deep design §39/§40을 독립 재검토해 selected/implemented-candidate/not-implemented를 실제
|
||
evidence와 일치시킨다.
|
||
- Sentinel/Cluster/k3s/R3 evidence가 없으면 지원/완료로 표기하지 않는다.
|
||
- Task 12의 전체 검증을 실행하고 동시 작업의 비-Redis 실패는 소유 파일과 증거를 분리한다.
|
||
- Redis README/spec/runbook, readiness registry, CI artifact 계약을 동기화한다.
|
||
- LLM Wiki branch-note와 실제 파생 raw 문서를 양방향 링크로 캡처한다.
|