18 KiB
title, source_type, track, status, status_label, difficulty, duration_estimate, prerequisites, parent_project, parent_branch, target_date, created, tags
| title | source_type | track | status | status_label | difficulty | duration_estimate | prerequisites | parent_project | parent_branch | target_date | created | tags | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| daily-task / infra / actuator-readiness-probe-db-disconnect-detection | daily-task | infra | raw | not-started | intermediate | 120 |
|
ca-skeleton-operational-contract | 2026-05-29 | 2026-05-28 |
|
daily-task / infra / actuator-readiness-probe-db-disconnect-detection
Layer:
raw/daily-tasks/infra/— 인프라/운영 트랙 일일 실습 과제.status_label:not-started→in-progress→donedifficulty:intermediate(Spring Boot actuator 기본 사용 + k8s probe 개념 가정)duration_estimate: 120 (Apply / 측정 대기 시간 포함)이 과제의 위치: infra 트랙 1일차. raw/project-notes/ca-skeleton-operational-contract §15 (Runtime/Lifecycle) — actuator health/readiness/liveness 기준 — 의 측정 가능한 1차 검증. develop 첫 과제 (
archunit-controller-domain-return-rule) 와 같은 날 진행해 코드 contract + 운영 contract 가 한 사이클에 검증되는 경험을 만든다.
Parent / 부모 (필수)
- Parent project: raw/project-notes/ca-skeleton-operational-contract (§15 Runtime/Lifecycle, §18 Metrics/Alerting)
- 연관 branch: (없음 — operational contract 직접 검증)
1. 학습 목표 / Learning Objectives
- L1: Spring Boot
health/readiness와health/liveness의 의미 차이 — 내 서비스가 트래픽 받을 준비됐는가 (readiness) vs 프로세스를 죽여야 하는가 (liveness) — 를 1분 안에 누군가에게 설명할 수 있다 - L2:
application.yaml에 actuator health group 을 명시 설정하고/actuator/health/readiness에 DB indicator 가 포함됨을 검증할 수 있다 - L3: DB 단절 시 readiness 가
OUT_OF_SERVICE로 전환되고 이 변화가 몇 초 만에 (kubectl + prometheus 양 채널) 표면화되는지 측정값으로 제시할 수 있다 - L4 (필수, advanced): liveness 는 동일 상황에서 전환되지 않음 (pod kill ≠ DB 단절) 을 확인하고, 왜 그래야 하는지 trade-off 로 설명할 수 있다 — 이 한 줄이 mid 와 senior 의 차이
2. 스토리라인 / WHY (Storyline)
raw/project-notes/ca-skeleton-operational-contract §15 는 "actuator health/readiness/liveness 기준" 을 요구하지만, 많은 프로젝트가 default /actuator/health 만 보는 readinessProbe 로 만족한다. 이 default 의 의미는 "프로세스가 살아있다" 이지 "트래픽 받을 준비됐다" 가 아니다.
DB가 죽어도 Spring Boot 프로세스는 잘 살아있으니 /actuator/health 는 200을 반환하고, k8s readinessProbe 는 ready 라고 판정하고, 트래픽이 흘러오고, 5xx 가 양산된다. 알림이 울리고 사람이 새벽에 깨고, root cause 는 "왜 우리는 DB 단절을 readiness 에 반영하지 않았는가" 가 된다.
오늘은 그 한 가지 — readiness 를 명시적으로 분리하고 DB indicator 를 포함 — 를 설정하고, 의도적으로 DB 를 끊었을 때 몇 초 후 not-ready 가 어디서 어떻게 표면화되는지 측정값으로 답할 수 있게 만든다.
심화 (L4): liveness 는 같은 상황에서 전환되지 않아야 한다. 왜냐하면 DB 단절은 프로세스를 죽일 이유 가 아니라 트래픽을 잠시 차단할 이유 이기 때문. 이걸 헷갈리면 pod 이 재기동 폭주에 들어가서 DB 가 살아나도 cluster 가 회복 불능. 이 trade-off 가 시니어 초반급 사고의 핵심.
3. 환경 / Environment
작업 호스트: 로컬 Linux/macOS/WSL2 (사용자 환경에 맞게)
대상 환경:
- Cluster: 로컬
kind또는k3d(cluster 없으면 시작 절차에 포함) - Namespace:
ca-tmpl-dev - Kubeconfig context:
kind-ca-tmpl-dev(예시)
도구 버전:
kubectl: 1.30+kind: 0.23+ (또는k3d5.6+, 또는 minikube)docker: 24.x- Spring Boot: 3.3.x (ca-tmpl 기존)
- 관측: Prometheus 2.50+ + Grafana 11.x (kube-prometheus-stack helm chart 권장)
사전 셋업:
# 1) 작업 디렉토리 + 브랜치
cd ~/workspace/ca-tmpl-infra # (또는 ca-tmpl 의 deploy/ 디렉토리)
git checkout -b daily-task/infra/actuator-readiness-probe-db-disconnect-detection
# 2) cluster 확인
kubectl config current-context
kubectl get ns ca-tmpl-dev || kubectl create ns ca-tmpl-dev
# 3) 현재 상태 스냅샷 (롤백 reference)
kubectl get all -n ca-tmpl-dev -o yaml > /tmp/snapshot-pre-readiness-probe.yaml
# 4) Prometheus / Grafana 준비 (없으면 설치)
helm list -n monitoring | grep prometheus || echo "kube-prometheus-stack 설치 필요"
# 5) 현재 ca-tmpl 의 application.yaml 확인
grep -A 10 'management:' ca-tmpl/src/main/resources/application.yaml || echo "actuator 설정 없음"
변경 예정 리소스:
ca-tmpl/src/main/resources/application.yaml—management.endpoint.health.probes.enabled=true, group readiness/liveness 명시deploy/k8s/ca-tmpl-deployment.yaml— readinessProbe path 분리, livenessProbe 의 thresholds 명시- (선택)
deploy/k8s/alerts/db-disconnect.yaml— PrometheusRule 신규
4. 사전 지식 / Prerequisites
- raw/project-notes/ca-skeleton-operational-contract §15 (Runtime/Lifecycle) + §18 (Metrics/Alerting) 정독
- raw/official-docs/runtime-health-spring-actuator-groups — actuator health group 공식 spec
- (있으면) Kubernetes liveness vs readiness 공식 정의 —
kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/ - Spring Boot
DataSourceHealthIndicator의 default 동작 (connection validation query)
5. 단계별 과제 / Exercises
Step 1: 베이스라인 측정 (~20min)
- What: 현재 상태를 수치 로 기록. 변경 후 비교 가능해야 함.
- How (hint):
- 현재
/actuator/health응답 body (DB indicator 가 있는지 없는지) kubectl describe pod <ca-tmpl-pod>→ readinessProbe / livenessProbe 설정 (path, initialDelay, period, threshold)kubectl get pod -w로 ready 상태 watch- DB container 가 살아있는 동안의 readiness 응답 시간 (curl -w 로 측정)
- 현재
- Done when: §7 결과물 섹션에 baseline 표 3행 이상 (
/actuator/health응답 type / readinessProbe path / readiness latency)
Step 2: actuator group 설정 + manifest 변경 (~30min)
- What:
application.yaml에 health group 명시, k8s manifest 의 probe path 분리. - How (hint):
# application.yaml
management:
endpoint:
health:
probes:
enabled: true
group:
readiness:
include: readinessState,db,diskSpace
liveness:
include: livenessState
show-details: never # PII / secret leak 방지 (CLAUDE.md §11)
# k8s deployment.yaml (발췌)
spec:
containers:
- name: ca-tmpl
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 3 # = 15초 후 NotReady
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 6 # = 60초 후 kill (보수적)
- 함정 / 의도적 노출:
- readiness 에
db를 너무 빨리 포함시키면 부팅 시점에 DB 가 천천히 ready 되는 동안 pod 도 NotReady → 부팅 지연 - liveness 에
db를 포함시키면 DB 죽었다고 pod kill — 이게 가장 큰 함정. 의도적으로 절대 안 한다.
- readiness 에
- Done when:
kubectl apply --dry-run=server -f <manifest>통과- probe path / period / threshold 가 baseline 과 어떻게 다른지 diff 검토 완료
Step 3: Apply + 정상 readiness 확인 (~25min)
- What: 실제 apply, rollout 대기, 정상 상태 측정.
- How (hint):
kubectl apply -f deploy/k8s/ca-tmpl-deployment.yaml
kubectl rollout status deployment/ca-tmpl -n ca-tmpl-dev --timeout=120s
# 1) HTTP 응답 직접 확인
kubectl port-forward svc/ca-tmpl 8080:8080 -n ca-tmpl-dev &
curl -sS http://localhost:8080/actuator/health/readiness | jq .
curl -sS http://localhost:8080/actuator/health/liveness | jq .
# 2) k8s pod 상태
kubectl get pod -n ca-tmpl-dev -l app=ca-tmpl
# 3) prometheus query (kube-state-metrics)
# promql: kube_pod_container_status_ready{namespace="ca-tmpl-dev",container="ca-tmpl"}
- Done when:
- readiness 응답 =
{"status":"UP"}(show-details=never 로 detail 미노출 — §11 정합) - kubectl
READY 1/1 - prometheus 의
kube_pod_container_status_ready= 1
- readiness 응답 =
Step 4: 의도적 DB 단절 → not-ready 전환 시간 측정 (~25min, 본 과제의 핵심)
- What: DB 를 끊고 몇 초 후 readiness 가 false 로 전환되는지 4-5 채널 교차 측정. liveness 는 전환되지 않음 을 확인.
- How (hint):
# 1) 측정 시작 시각 기록
TS_START=$(date +%s)
echo "DB cut at $TS_START"
# 2) DB 단절 (postgres container stop 또는 service block)
kubectl delete pod -n ca-tmpl-dev -l app=postgres
# (또는) docker stop ca-tmpl-postgres
# 3) 즉시 watch 시작 — 별 터미널에서:
watch -n 1 "kubectl get pod -n ca-tmpl-dev -l app=ca-tmpl -o wide && curl -sS http://localhost:8080/actuator/health/readiness; echo; curl -sS http://localhost:8080/actuator/health/liveness"
# 4) NotReady 표면화 시각 측정
# - readiness 응답이 503 또는 OUT_OF_SERVICE 로 바뀌는 순간
# - kubectl 의 READY 가 0/1 로 바뀌는 순간
# - prometheus 의 metric 이 0 으로 바뀌는 순간
# 세 값의 차이 자체가 학습 포인트
# 5) liveness 가 *전환되지 않는지* 확인 (UP 유지)
-
측정해야 할 값들:
- T_actuator: actuator readiness 가 OUT_OF_SERVICE 로 전환된 시각 (DB 단절 후 N초)
- T_kubectl:
kubectl get pod의 READY 가 0/1 로 표시되는 시각 - T_prometheus: prometheus metric 이 0 으로 바뀌는 시각 (kube-state-metrics scrape interval 의 영향)
- liveness 응답 상태: 반드시 UP 유지
-
함정 / 트레이드오프 의식 (시니어 사고):
failureThreshold=3,periodSeconds=5이면 최대 15초 후 표면화 — 더 빨리 잡으려면 period↓ 인데 false positive ↑- HikariCP 의
connection-timeout과 actuator probe timeout 의 상호작용 — actuator가 DB indicator 평가 시 30초 hang 하면 readinessProbe 자체도 timeout - prometheus scrape interval (예: 30초) 이 alert 표면화의 lower bound — 5초마다 NotReady 가 토글되면 prometheus 는 못 봄. 이걸 모르면 "왜 alert 가 안 울리지" 미스터리 발생.
-
Done when:
- 세 측정값 (T_actuator, T_kubectl, T_prometheus) 표로 기록
- liveness 가 전환되지 않음 명시적으로 확인
- §8 회고에 "왜 세 값이 다른가" 한 문장 답변
Step 5 (선택, advanced): DB 복원 → readiness 자동 복귀 측정 (~20min)
- What: DB 다시 살리고 readiness 가 자동으로 UP 으로 돌아오는 시간 측정 + 그 사이 traffic 처리 동작 확인.
- How (hint):
- DB pod 재시작
- HikariCP 의 connection pool 이 자동 복구되는지 (
hikari.minimum-idle영향) - 복귀 시간 = HikariCP retry interval + actuator probe period
- 트레이드오프 (시니어 사고):
- 자동 복구가 너무 빠르면 DB 가 flaky 할 때 readiness 가 토글 — load balancer 도 토글
- 의도적 hysteresis 권장 (예: 30초 연속 UP 일 때만 ready)
- Done when: 복귀 시간 측정값 + 그 사이 in-flight 요청의 운명 (drop / 502 / queue) 기록
6. 검증 / Assessment
자동 검증 (4-5 채널 중 ≥2개 교차):
# 1) Probe / health (정상 상태)
curl -fsS http://localhost:8080/actuator/health/readiness | jq -e '.status == "UP"'
curl -fsS http://localhost:8080/actuator/health/liveness | jq -e '.status == "UP"'
# 합격 기준: 두 명령 모두 exit 0
# 2) k8s 리소스 상태 (rollout 후)
kubectl rollout status deployment/ca-tmpl -n ca-tmpl-dev --timeout=60s
# 합격 기준: successfully rolled out
# 3) PromQL — readiness 가 metric 으로 노출
# 권장 query: kube_pod_container_status_ready{namespace="ca-tmpl-dev",container="ca-tmpl"}
# 합격 기준: 정상 시 = 1
# 4) DB 단절 시뮬레이션 시 readiness 전환
# (Step 4 의 측정 결과를 contract test 로 만들 수 있다면 가산점)
# 5) Smoke test — 정상 endpoint 가 200 응답
curl -fsS http://localhost:8080/api/v1/<sample-endpoint>
# 합격 기준: exit 0 (정상 상태에서)
수동 self-check:
- 위 4-5채널 중 ≥2 가 교차 확인됨 (단일 채널 의존 금지)
- DB 단절 시 readiness 전환 시간이 measurable (Step 4 측정값 표 존재)
- liveness 가 DB 단절 상황에서 UP 유지 — 측정으로 확인
- 롤백 명령 (
kubectl apply -f /tmp/snapshot-pre-readiness-probe.yaml) 이 완전히 베이스라인으로 복귀 가능 - L1~L4 학습 목표 모두 수행 가능 — 특히 L4 (liveness/readiness trade-off) 를 한 줄로 설명 가능
- manifest commit 메시지가 "왜" 를 답함
7. 결과물 / Outcomes
- commit / PR:
- 브랜치:
daily-task/infra/actuator-readiness-probe-db-disconnect-detection - commits: <해시 + 1줄>
- 브랜치:
- 변경된 manifest / 설정:
ca-tmpl/src/main/resources/application.yaml— actuator health group 명시deploy/k8s/ca-tmpl-deployment.yaml— probe path 분리, threshold 명시
- 측정값 표 (Step 1 baseline vs Step 4 적용 후):
| 측정 항목 | Baseline | DB 단절 후 |
|---|---|---|
/actuator/health/readiness 응답 |
UP / 200 | OUT_OF_SERVICE / 503 (T초 후) |
kubectl get pod READY |
1/1 | 0/1 (T초 후) |
prometheus kube_pod_container_status_ready |
1 | 0 (T초 후) |
| liveness 응답 | UP | UP 유지 (의도) |
- Dashboard / Alert:
- Grafana panel:
ca-tmpl readiness(kube_pod_container_status_ready over time) - Alert rule (작성 시): readiness=0 이 60초 지속 시 P2 alert
- Grafana panel:
- Runbook stub:
- 알람 발생 시 1차 확인:
kubectl describe pod -l app=ca-tmpl+curl /actuator/health/readiness - 즉시 fail-fast vs degrade: DB 단절 = readiness 차단 (fail-fast), pod kill 아님 (degrade with traffic block)
- 알람 발생 시 1차 확인:
- 학습한 개념 (wiki/concepts 후보):
- readiness vs liveness 의 운영적 차이
- HikariCP connection timeout 과 probe timeout 의 상호작용
- prometheus scrape interval 이 alert detection 의 lower bound
- 다음 과제 thread:
- HikariCP
connection-timeout의 적정값 측정 - readinessProbe failure 후 traffic drain (Kubernetes service endpoint 갱신 시간)
- chaos test 자동화 (chaos-mesh)
- circuit breaker (Resilience4j) 와 readiness 의 관계
- HikariCP
8. 회고 / Reflection (~5min)
- 막혔던 곳 (몇 분 / 어디서):
- 예상과 다른 점 (특히 측정값 vs 예측):
- 예:
failureThreshold=3인데 readiness 가 15초보다 늦게 표면화 — 왜? (probe timeout? actuator hang?) - prometheus metric 이 훨씬 늦게 변함 — scrape interval 영향
- 예:
- 다음 반복에서 개선할 점:
- 부수 효과로 발견한 것:
- 이 과제의 난이도가 적정했는가:
너무 쉬움/적정/너무 어려움 - 시니어 사고 체크 (필수 1줄 답변):
- "왜 liveness 에 DB 를 포함하면 안 되는가?" — <답>
- "T_actuator, T_kubectl, T_prometheus 세 값이 다른 이유는 무엇인가?" — <답>
- "readiness 토글 (UP→OUT_OF_SERVICE→UP) 이 잦으면 어떤 운영 문제를 일으키는가?" — <답>
9. 출처 / Sources
| Source | 정당화 영역 |
|---|---|
| raw/company-tech-blogs/skillable-hands-on-lab-structure | template 9-section 구조 |
| raw/company-tech-blogs/deliberate-practice-software-developers-redgreencode | §5 단계 분할 + §8 reflection |
| raw/project-notes/ca-skeleton-operational-contract | §15 Runtime/Lifecycle (probe 기준) + §18 Metrics/Alerting |
| raw/official-docs/runtime-health-spring-actuator-groups | actuator health group 공식 spec — readiness/liveness 분리 근거 |
10. 완료 후 정리 / Closure
- 최종 status_label:
done|abandoned - 소요 시간 실측: <분> (vs 120) — 차이는 §8 회고에
- promotable 후보:
actually-implemented→ ca-skeleton-operational-contract §15 의 actuator probe 분리 결정의 실 구현 증거locally-verified→ DB 단절 → readiness 전환 측정값 4채널 교차 확인prod-verified→ (해당 없음 — 로컬 cluster)
- 추출하지 않을 항목:
- chaos-mesh 자동화 / circuit breaker 통합 — 별도 daily-task 로 분할
11. 운영 회복력 / Operational Resilience (infra 전용 anchor)
- 본 변경이 도입하는 새 실패 모드:
- DB indicator 가 시간이 오래 걸리는 query 면 readinessProbe 자체가 timeout → false NotReady
- probe period 가 너무 짧으면 DB 가 잠시 hiccup 할 때 ready 토글 → load balancer 토글 → 502 spike
- 새 실패 모드의 fail-fast vs degrade 분류:
- DB 단절 = fail-fast (트래픽 차단)
- DB 응답 지연 = degrade (slow 응답이지만 트래픽 유지) — readiness 에 포함시킬지 결정 필요
- 모니터링 누락 위험:
- prometheus scrape interval 보다 짧은 not-ready 윈도우는 못 봄 (false success)
- alert quiet hours 가 없으면 readiness toggle 시 alert 폭주
- 롤백 트리거 조건:
- readiness false 가 5분 지속 + DB 자체는 정상 → 본 변경 자체의 false positive 가능성 → 즉시 롤백
kubectl apply -f /tmp/snapshot-pre-readiness-probe.yaml
- 연관 alert / runbook:
- raw/project-notes/ca-skeleton-operational-contract §28 Operational Runbook 의 "DB unavailable" 시나리오와 정합
- 본 과제의 PrometheusRule 이 §28 의 1차 alert 항목으로 등록되어야 함