init: llm-wiki-haness 하네스 설계

This commit is contained in:
DongHyeonka
2026-07-24 14:21:35 +09:00
parent 42bf3db4fd
commit 6c53ded9cb
2436 changed files with 194486 additions and 1 deletions
@@ -0,0 +1,360 @@
---
title: daily-task / infra / actuator-readiness-probe-db-disconnect-detection
source_type: daily-task
track: infra
status: raw
status_label: not-started
difficulty: intermediate
duration_estimate: 120
prerequisites:
- "[[raw/project-notes/ca-skeleton-operational-contract]]"
- "[[raw/official-docs/runtime-health-spring-actuator-groups]]"
parent_project: ca-skeleton-operational-contract
parent_branch:
target_date: 2026-05-29
created: 2026-05-28
tags: [daily-task, infra, observability, runtime]
---
# daily-task / infra / actuator-readiness-probe-db-disconnect-detection
> Layer: `raw/daily-tasks/infra/` — **인프라/운영 트랙 일일 실습 과제**.
> `status_label`: `not-started` → `in-progress` → `done`
> `difficulty`: `intermediate` (Spring Boot actuator 기본 사용 + k8s probe 개념 가정)
> `duration_estimate`: 120 (Apply / 측정 대기 시간 포함)
>
> **이 과제의 위치**: infra 트랙 1일차. [[raw/project-notes/ca-skeleton-operational-contract]] §15 (Runtime/Lifecycle) — actuator health/readiness/liveness 기준 — 의 *측정 가능한 1차 검증*. develop 첫 과제 (`archunit-controller-domain-return-rule`) 와 같은 날 진행해 코드 contract + 운영 contract 가 한 사이클에 검증되는 경험을 만든다.
## Parent / 부모 (필수)
- **Parent project**: [[raw/project-notes/ca-skeleton-operational-contract]] (§15 Runtime/Lifecycle, §18 Metrics/Alerting)
- **연관 branch**: (없음 — operational contract 직접 검증)
## 1. 학습 목표 / Learning Objectives
- [ ] **L1**: Spring Boot `health/readiness``health/liveness` 의 의미 차이 — *내 서비스가 트래픽 받을 준비됐는가* (readiness) vs *프로세스를 죽여야 하는가* (liveness) — 를 1분 안에 누군가에게 설명할 수 있다
- [ ] **L2**: `application.yaml` 에 actuator health group 을 명시 설정하고 `/actuator/health/readiness` 에 DB indicator 가 포함됨을 검증할 수 있다
- [ ] **L3**: DB 단절 시 readiness 가 `OUT_OF_SERVICE` 로 전환되고 이 변화가 *몇 초 만에* (kubectl + prometheus 양 채널) 표면화되는지 *측정값으로* 제시할 수 있다
- [ ] **L4 (필수, advanced)**: liveness 는 *동일 상황에서 전환되지 않음* (pod kill ≠ DB 단절) 을 확인하고, 왜 그래야 하는지 trade-off 로 설명할 수 있다 — 이 한 줄이 mid 와 senior 의 차이
## 2. 스토리라인 / WHY (Storyline)
[[raw/project-notes/ca-skeleton-operational-contract]] §15 는 "actuator health/readiness/liveness 기준" 을 요구하지만, 많은 프로젝트가 default `/actuator/health` 만 보는 readinessProbe 로 만족한다. 이 default 의 의미는 **"프로세스가 살아있다"** 이지 **"트래픽 받을 준비됐다"** 가 아니다.
DB가 죽어도 Spring Boot 프로세스는 잘 살아있으니 `/actuator/health` 는 200을 반환하고, k8s readinessProbe 는 *ready* 라고 판정하고, 트래픽이 흘러오고, 5xx 가 양산된다. 알림이 울리고 사람이 새벽에 깨고, root cause 는 "왜 우리는 DB 단절을 readiness 에 반영하지 않았는가" 가 된다.
오늘은 *그 한 가지* — readiness 를 명시적으로 분리하고 DB indicator 를 포함 — 를 설정하고, 의도적으로 DB 를 *끊었을 때* 몇 초 후 not-ready 가 어디서 어떻게 표면화되는지 *측정값으로* 답할 수 있게 만든다.
심화 (L4): liveness 는 같은 상황에서 *전환되지 않아야* 한다. 왜냐하면 DB 단절은 *프로세스를 죽일 이유* 가 아니라 *트래픽을 잠시 차단할 이유* 이기 때문. 이걸 헷갈리면 pod 이 재기동 폭주에 들어가서 DB 가 살아나도 cluster 가 회복 불능. 이 trade-off 가 시니어 초반급 사고의 핵심.
## 3. 환경 / Environment
**작업 호스트**: 로컬 Linux/macOS/WSL2 (사용자 환경에 맞게)
**대상 환경**:
- Cluster: 로컬 `kind` 또는 `k3d` (cluster 없으면 시작 절차에 포함)
- Namespace: `ca-tmpl-dev`
- Kubeconfig context: `kind-ca-tmpl-dev` (예시)
**도구 버전**:
- `kubectl`: 1.30+
- `kind`: 0.23+ (또는 `k3d` 5.6+, 또는 minikube)
- `docker`: 24.x
- Spring Boot: 3.3.x (ca-tmpl 기존)
- 관측: Prometheus 2.50+ + Grafana 11.x (kube-prometheus-stack helm chart 권장)
**사전 셋업**:
```bash
# 1) 작업 디렉토리 + 브랜치
cd ~/workspace/ca-tmpl-infra # (또는 ca-tmpl 의 deploy/ 디렉토리)
git checkout -b daily-task/infra/actuator-readiness-probe-db-disconnect-detection
# 2) cluster 확인
kubectl config current-context
kubectl get ns ca-tmpl-dev || kubectl create ns ca-tmpl-dev
# 3) 현재 상태 스냅샷 (롤백 reference)
kubectl get all -n ca-tmpl-dev -o yaml > /tmp/snapshot-pre-readiness-probe.yaml
# 4) Prometheus / Grafana 준비 (없으면 설치)
helm list -n monitoring | grep prometheus || echo "kube-prometheus-stack 설치 필요"
# 5) 현재 ca-tmpl 의 application.yaml 확인
grep -A 10 'management:' ca-tmpl/src/main/resources/application.yaml || echo "actuator 설정 없음"
```
**변경 예정 리소스**:
- `ca-tmpl/src/main/resources/application.yaml``management.endpoint.health.probes.enabled=true`, group readiness/liveness 명시
- `deploy/k8s/ca-tmpl-deployment.yaml` — readinessProbe path 분리, livenessProbe 의 thresholds 명시
- (선택) `deploy/k8s/alerts/db-disconnect.yaml` — PrometheusRule 신규
## 4. 사전 지식 / Prerequisites
- [[raw/project-notes/ca-skeleton-operational-contract]] §15 (Runtime/Lifecycle) + §18 (Metrics/Alerting) 정독
- [[raw/official-docs/runtime-health-spring-actuator-groups]] — actuator health group 공식 spec
- (있으면) Kubernetes liveness vs readiness 공식 정의 — `kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/`
- Spring Boot `DataSourceHealthIndicator` 의 default 동작 (connection validation query)
## 5. 단계별 과제 / Exercises
### Step 1: 베이스라인 측정 (~20min)
- **What**: 현재 상태를 *수치* 로 기록. 변경 후 비교 가능해야 함.
- **How (hint)**:
- 현재 `/actuator/health` 응답 body (DB indicator 가 *있는지* 없는지)
- `kubectl describe pod <ca-tmpl-pod>` → readinessProbe / livenessProbe 설정 (path, initialDelay, period, threshold)
- `kubectl get pod -w` 로 ready 상태 watch
- DB container 가 살아있는 동안의 readiness 응답 시간 (curl -w 로 측정)
- **Done when**: §7 결과물 섹션에 baseline 표 3행 이상 (`/actuator/health` 응답 type / readinessProbe path / readiness latency)
### Step 2: actuator group 설정 + manifest 변경 (~30min)
- **What**: `application.yaml` 에 health group 명시, k8s manifest 의 probe path 분리.
- **How (hint)**:
```yaml
# application.yaml
management:
endpoint:
health:
probes:
enabled: true
group:
readiness:
include: readinessState,db,diskSpace
liveness:
include: livenessState
show-details: never # PII / secret leak 방지 (CLAUDE.md §11)
```
```yaml
# k8s deployment.yaml (발췌)
spec:
containers:
- name: ca-tmpl
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 3 # = 15초 후 NotReady
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 6 # = 60초 후 kill (보수적)
```
- **함정 / 의도적 노출**:
- readiness 에 `db`*너무 빨리* 포함시키면 부팅 시점에 DB 가 천천히 ready 되는 동안 pod 도 NotReady → 부팅 지연
- liveness 에 `db` 를 포함시키면 *DB 죽었다고 pod kill***이게 가장 큰 함정. 의도적으로 절대 안 한다.**
- **Done when**:
- `kubectl apply --dry-run=server -f <manifest>` 통과
- probe path / period / threshold 가 baseline 과 어떻게 다른지 diff 검토 완료
### Step 3: Apply + 정상 readiness 확인 (~25min)
- **What**: 실제 apply, rollout 대기, 정상 상태 측정.
- **How (hint)**:
```bash
kubectl apply -f deploy/k8s/ca-tmpl-deployment.yaml
kubectl rollout status deployment/ca-tmpl -n ca-tmpl-dev --timeout=120s
# 1) HTTP 응답 직접 확인
kubectl port-forward svc/ca-tmpl 8080:8080 -n ca-tmpl-dev &
curl -sS http://localhost:8080/actuator/health/readiness | jq .
curl -sS http://localhost:8080/actuator/health/liveness | jq .
# 2) k8s pod 상태
kubectl get pod -n ca-tmpl-dev -l app=ca-tmpl
# 3) prometheus query (kube-state-metrics)
# promql: kube_pod_container_status_ready{namespace="ca-tmpl-dev",container="ca-tmpl"}
```
- **Done when**:
- readiness 응답 = `{"status":"UP"}` (show-details=never 로 detail 미노출 — §11 정합)
- kubectl `READY 1/1`
- prometheus 의 `kube_pod_container_status_ready` = 1
### Step 4: 의도적 DB 단절 → not-ready 전환 시간 측정 (~25min, **본 과제의 핵심**)
- **What**: DB 를 *끊고* 몇 초 후 readiness 가 false 로 전환되는지 4-5 채널 교차 측정. liveness 는 전환되지 *않음* 을 확인.
- **How (hint)**:
```bash
# 1) 측정 시작 시각 기록
TS_START=$(date +%s)
echo "DB cut at $TS_START"
# 2) DB 단절 (postgres container stop 또는 service block)
kubectl delete pod -n ca-tmpl-dev -l app=postgres
# (또는) docker stop ca-tmpl-postgres
# 3) 즉시 watch 시작 — 별 터미널에서:
watch -n 1 "kubectl get pod -n ca-tmpl-dev -l app=ca-tmpl -o wide && curl -sS http://localhost:8080/actuator/health/readiness; echo; curl -sS http://localhost:8080/actuator/health/liveness"
# 4) NotReady 표면화 시각 측정
# - readiness 응답이 503 또는 OUT_OF_SERVICE 로 바뀌는 순간
# - kubectl 의 READY 가 0/1 로 바뀌는 순간
# - prometheus 의 metric 이 0 으로 바뀌는 순간
# 세 값의 차이 자체가 학습 포인트
# 5) liveness 가 *전환되지 않는지* 확인 (UP 유지)
```
- **측정해야 할 값들**:
- T_actuator: actuator readiness 가 OUT_OF_SERVICE 로 전환된 시각 (DB 단절 후 N초)
- T_kubectl: `kubectl get pod` 의 READY 가 0/1 로 표시되는 시각
- T_prometheus: prometheus metric 이 0 으로 바뀌는 시각 (kube-state-metrics scrape interval 의 영향)
- liveness 응답 상태: *반드시* UP 유지
- **함정 / 트레이드오프 의식** (시니어 사고):
- `failureThreshold=3`, `periodSeconds=5` 이면 *최대* 15초 후 표면화 — 더 빨리 잡으려면 period↓ 인데 false positive ↑
- HikariCP 의 `connection-timeout` 과 actuator probe timeout 의 상호작용 — actuator가 DB indicator 평가 시 30초 hang 하면 readinessProbe 자체도 timeout
- **prometheus scrape interval (예: 30초) 이 alert 표면화의 lower bound** — 5초마다 NotReady 가 토글되면 prometheus 는 못 봄. 이걸 모르면 "왜 alert 가 안 울리지" 미스터리 발생.
- **Done when**:
- 세 측정값 (T_actuator, T_kubectl, T_prometheus) 표로 기록
- liveness 가 *전환되지 않음* 명시적으로 확인
- **§8 회고에 "왜 세 값이 다른가" 한 문장 답변**
### Step 5 (선택, advanced): DB 복원 → readiness 자동 복귀 측정 (~20min)
- **What**: DB 다시 살리고 readiness 가 자동으로 UP 으로 돌아오는 시간 측정 + 그 사이 traffic 처리 동작 확인.
- **How (hint)**:
- DB pod 재시작
- HikariCP 의 connection pool 이 자동 복구되는지 (`hikari.minimum-idle` 영향)
- 복귀 시간 = HikariCP retry interval + actuator probe period
- **트레이드오프** (시니어 사고):
- 자동 복구가 *너무 빠르면* DB 가 flaky 할 때 readiness 가 토글 — load balancer 도 토글
- 의도적 hysteresis 권장 (예: 30초 연속 UP 일 때만 ready)
- **Done when**: 복귀 시간 측정값 + 그 사이 in-flight 요청의 운명 (drop / 502 / queue) 기록
## 6. 검증 / Assessment
**자동 검증** (4-5 채널 중 ≥2개 교차):
```bash
# 1) Probe / health (정상 상태)
curl -fsS http://localhost:8080/actuator/health/readiness | jq -e '.status == "UP"'
curl -fsS http://localhost:8080/actuator/health/liveness | jq -e '.status == "UP"'
# 합격 기준: 두 명령 모두 exit 0
# 2) k8s 리소스 상태 (rollout 후)
kubectl rollout status deployment/ca-tmpl -n ca-tmpl-dev --timeout=60s
# 합격 기준: successfully rolled out
# 3) PromQL — readiness 가 metric 으로 노출
# 권장 query: kube_pod_container_status_ready{namespace="ca-tmpl-dev",container="ca-tmpl"}
# 합격 기준: 정상 시 = 1
# 4) DB 단절 시뮬레이션 시 readiness 전환
# (Step 4 의 측정 결과를 contract test 로 만들 수 있다면 가산점)
# 5) Smoke test — 정상 endpoint 가 200 응답
curl -fsS http://localhost:8080/api/v1/<sample-endpoint>
# 합격 기준: exit 0 (정상 상태에서)
```
**수동 self-check**:
- [ ] 위 4-5채널 중 ≥2 가 *교차* 확인됨 (단일 채널 의존 금지)
- [ ] DB 단절 시 readiness 전환 시간이 measurable (Step 4 측정값 표 존재)
- [ ] liveness 가 DB 단절 상황에서 *UP 유지* — 측정으로 확인
- [ ] 롤백 명령 (`kubectl apply -f /tmp/snapshot-pre-readiness-probe.yaml`) 이 *완전히* 베이스라인으로 복귀 가능
- [ ] L1~L4 학습 목표 모두 *수행 가능* — 특히 L4 (liveness/readiness trade-off) 를 *한 줄로* 설명 가능
- [ ] manifest commit 메시지가 "왜" 를 답함
## 7. 결과물 / Outcomes
- **commit / PR**:
- 브랜치: `daily-task/infra/actuator-readiness-probe-db-disconnect-detection`
- commits: <해시 + 1줄>
- **변경된 manifest / 설정**:
- `ca-tmpl/src/main/resources/application.yaml` — actuator health group 명시
- `deploy/k8s/ca-tmpl-deployment.yaml` — probe path 분리, threshold 명시
- **측정값 표** (Step 1 baseline vs Step 4 적용 후):
| 측정 항목 | Baseline | DB 단절 후 |
|---|---|---|
| `/actuator/health/readiness` 응답 | UP / 200 | OUT_OF_SERVICE / 503 (T초 후) |
| `kubectl get pod` READY | 1/1 | 0/1 (T초 후) |
| prometheus `kube_pod_container_status_ready` | 1 | 0 (T초 후) |
| liveness 응답 | UP | **UP 유지** (의도) |
- **Dashboard / Alert**:
- Grafana panel: `ca-tmpl readiness` (kube_pod_container_status_ready over time)
- Alert rule (작성 시): readiness=0 이 60초 지속 시 P2 alert
- **Runbook stub**:
- 알람 발생 시 1차 확인: `kubectl describe pod -l app=ca-tmpl` + `curl /actuator/health/readiness`
- 즉시 fail-fast vs degrade: DB 단절 = readiness 차단 (fail-fast), pod kill 아님 (degrade with traffic block)
- **학습한 개념** (wiki/concepts 후보):
- readiness vs liveness 의 운영적 차이
- HikariCP connection timeout 과 probe timeout 의 상호작용
- prometheus scrape interval 이 alert detection 의 lower bound
- **다음 과제 thread**:
- HikariCP `connection-timeout` 의 적정값 측정
- readinessProbe failure 후 traffic drain (Kubernetes service endpoint 갱신 시간)
- chaos test 자동화 (chaos-mesh)
- circuit breaker (Resilience4j) 와 readiness 의 관계
## 8. 회고 / Reflection (~5min)
- **막혔던 곳** (몇 분 / 어디서):
- **예상과 다른 점** (특히 측정값 vs 예측):
- 예: `failureThreshold=3` 인데 readiness 가 *15초보다 늦게* 표면화 — 왜? (probe timeout? actuator hang?)
- prometheus metric 이 *훨씬 늦게* 변함 — scrape interval 영향
- **다음 반복에서 개선할 점**:
- **부수 효과로 발견한 것**:
- **이 과제의 난이도가 적정했는가**: `너무 쉬움` / `적정` / `너무 어려움`
- **시니어 사고 체크** (필수 1줄 답변):
- "왜 liveness 에 DB 를 포함하면 안 되는가?" — <답>
- "T_actuator, T_kubectl, T_prometheus 세 값이 다른 이유는 무엇인가?" — <답>
- "readiness 토글 (UP→OUT_OF_SERVICE→UP) 이 잦으면 어떤 운영 문제를 일으키는가?" — <답>
## 9. 출처 / Sources
| Source | 정당화 영역 |
|---|---|
| [[raw/company-tech-blogs/skillable-hands-on-lab-structure]] | template 9-section 구조 |
| [[raw/company-tech-blogs/deliberate-practice-software-developers-redgreencode]] | §5 단계 분할 + §8 reflection |
| [[raw/project-notes/ca-skeleton-operational-contract]] | §15 Runtime/Lifecycle (probe 기준) + §18 Metrics/Alerting |
| [[raw/official-docs/runtime-health-spring-actuator-groups]] | actuator health group 공식 spec — readiness/liveness 분리 근거 |
## 10. 완료 후 정리 / Closure
- **최종 status_label**: `done` | `abandoned`
- **소요 시간 실측**: <분> (vs 120) — 차이는 §8 회고에
- **promotable 후보**:
- `actually-implemented` → ca-skeleton-operational-contract §15 의 actuator probe 분리 결정의 *실 구현* 증거
- `locally-verified` → DB 단절 → readiness 전환 측정값 4채널 교차 확인
- `prod-verified` → (해당 없음 — 로컬 cluster)
- **추출하지 않을 항목**:
- chaos-mesh 자동화 / circuit breaker 통합 — 별도 daily-task 로 분할
## 11. 운영 회복력 / Operational Resilience (infra 전용 anchor)
- **본 변경이 도입하는 새 실패 모드**:
- DB indicator 가 *시간이 오래 걸리는 query* 면 readinessProbe 자체가 timeout → false NotReady
- probe period 가 *너무 짧으면* DB 가 잠시 hiccup 할 때 ready 토글 → load balancer 토글 → 502 spike
- **새 실패 모드의 fail-fast vs degrade 분류**:
- DB 단절 = fail-fast (트래픽 차단)
- DB 응답 지연 = degrade (slow 응답이지만 트래픽 유지) — readiness 에 포함시킬지 결정 필요
- **모니터링 누락 위험**:
- prometheus scrape interval 보다 *짧은* not-ready 윈도우는 못 봄 (false success)
- alert quiet hours 가 없으면 readiness toggle 시 alert 폭주
- **롤백 트리거 조건**:
- readiness false 가 5분 지속 + DB 자체는 정상 → 본 변경 자체의 false positive 가능성 → 즉시 롤백
- `kubectl apply -f /tmp/snapshot-pre-readiness-probe.yaml`
- **연관 alert / runbook**:
- [[raw/project-notes/ca-skeleton-operational-contract]] §28 Operational Runbook 의 "DB unavailable" 시나리오와 정합
- 본 과제의 PrometheusRule 이 §28 의 1차 alert 항목으로 등록되어야 함