Files
keycloak-pattern/docs/experiment-a8-rolling-restart.md
T
DongHyeonkaandClaude Opus 5 e0d27d47ce docs: correct the places where documents contradicted their own evidence
An independent audit found ten documents printing values their evidence files do not contain. C-1 printed a session count of 0 where the evidence says 4, C-2 printed a success readback for a command that exited 1, and A-1 credited the conntrack flush with a split that the timestamps attribute to a pod restart four seconds earlier.

Also measured wal_writer_delay, which A-3 had asserted as matching without ever querying it, relabelled the A-6 control that moved 41 percent, noted A-8's nine-sample resolution, corrected D-1's RTO to the 41 seconds its own timeline shows, and added a correction banner to D-2. Every experiment document now links its evidence files with their real collection times, and the duplicate screenshots are documented as duplicates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 16:35:49 +09:00

223 lines
8.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# A-8 — 배포할 때마다 로그아웃되는가
브랜치 `feature/keycloak-a8-rolling-restart` ·
증거 [`docs/evidence/a8-rolling-restart/`](evidence/a8-rolling-restart/) ·
2026-09-04 13:3813:41 KST
**운영에서 가장 자주 겪는 일이다.** 장애가 아니라 정상 작업인데도
사용자가 로그아웃되면 그건 사고다.
---
## 0. 결론부터
| 확인 | 결과 |
|---|---|
| 재시작 중 서비스 중단 | **없음.** 전 구간 `200` |
| 재시작 전 발급한 refresh token | **여전히 통한다** (`200`) |
| DB 세션 수 | **151 → 151** 그대로 |
| 세션 캐시 | **0 으로 초기화** |
| 클러스터 | 자동 재형성 (`cluster_size = 2`) |
**세션은 살아남고 캐시만 사라진다.** 이것이 `persistent-user-sessions`
켜는 진짜 이유다.
---
## 1. 방법
```bash
# 1. 재시작 전 로그인 — 토큰을 상주 파드 안에 보관한다
kubectl -n keycloak-lab run a8-probe --image=curlimages/curl:8.11.1 \
--restart=Never --command -- sleep 3600
kubectl -n keycloak-lab exec a8-probe -- sh -c '<로그인 후 /tmp/rt, /tmp/sid 에 저장>'
# 2. 재시작하면서 5초 간격으로 외부 진입점을 찍는다
kubectl -n keycloak-lab rollout restart statefulset/keycloak
( for i in $(seq 1 48); do
curl -s -o /dev/null -w '%{http_code} ' --max-time 4 https://auth.hyeonworks.com/realms/master
sleep 5
done ) &
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=420s
```
**탐침 파드가 StatefulSet 밖에 있어야** 재시작을 넘어 토큰을 들고 있을 수 있다.
---
## 2. 가용성 — 무중단이었다
```
statefulset.apps/keycloak restarted
200 Waiting for partitioned roll out to finish: 0 out of 2 new pods have been updated...
Waiting for 1 pods to be ready...
200 200 200 200 Waiting for partitioned roll out to finish: 1 out of 2 new pods have been updated...
Waiting for 1 pods to be ready...
200 200 200 200 partitioned roll out complete: 2 new pods have been updated...
```
**9번 찍어서 9번 다 `200`.**
> **표본은 9개다.** 5초 간격으로 찍었으므로 **5초보다 짧은 끊김은 이 측정으로
> 잡히지 않는다.** 실제로 후속 작업에서 1초 간격·3초 타임아웃으로 재보니
> 롤백 전환 순간에 `000` 이 한 번 잡혔다
> ([`followup`](experiment-followup-untested-items.md) 2절).
> **"무중단" 은 관측 해상도에 달려 있으며, 여기서는 "5초 해상도에서 끊김이
> 관측되지 않았다" 까지가 정확한 서술이다.**
### 왜 무중단이 되는가
```
StatefulSet 롤링 재시작
├─ keycloak-1 종료 → Service 엔드포인트에서 빠짐
│ └─ 이 동안 keycloak-0 이 전부 받는다
├─ keycloak-1 기동 → readiness UP → 엔드포인트 복귀
└─ keycloak-0 종료 → ... (반복)
```
**한 번에 하나씩** 내리므로 항상 최소 하나는 Ready 다.
readiness 프로브가 이 전환을 정확히 맞춰준다 — A-2 에서 본 그 메커니즘이
여기서는 **정상 작업을 안전하게** 만든다.
> 다만 이 실험대는 **파드가 2개**다. replica 1 이면 반드시 끊긴다.
> 무중단은 공짜가 아니라 **replica ≥ 2 와 readiness 의 조합**이다.
---
## 3. 세션 생존
```
=== 재시작 전 발급한 refresh token 이 아직 통하는가 ===
대상 sid: XLcgQWRiJrTkuNZcJsNeT_2j
keycloak-0 에서 refresh HTTP 200
=== DB 에 그 세션이 남아 있는가 ===
user_session_id | created_on | last_session_refresh
--------------------------+------------+----------------------
XLcgQWRiJrTkuNZcJsNeT_2j | 1788495513 | 1788495577
전체 온라인 세션: 151 (재시작 전 151)
```
**`last_session_refresh``created_on` 보다 64초 뒤**다. 재시작 후의 refresh
가 **실제로 DB 에 기록**되었다는 뜻이다 — 응답 코드만 200 인 게 아니라
쓰기까지 정상이다.
**파드가 통째로 바뀌었는데(44초/66초 나이) 세션은 그대로다.**
---
## 4. 캐시는 사라진다
```
keycloak-0 sessions 캐시 0.0 건 / cluster_size 2.0
keycloak-1 sessions 캐시 1.0 건 / cluster_size 2.0
```
![캐시 초기화와 클러스터 재형성](evidence/a8-rolling-restart/a8-cache-reset-cluster-reformed.png)
**캐시는 프로세스 메모리이므로 재시작에 사라진다.** `keycloak-1` 의 1건은
방금 refresh 를 처리하며 새로 담은 것이다.
```
재시작 전: 캐시 N건 + DB 151건
재시작 후: 캐시 0건 + DB 151건 ← 진실은 DB 에 있다
```
**A-0 의 모델이 그대로 확인된다.** 캐시가 통째로 날아가도 정확성은 유지되고
**첫 접근만 느려진다** (룩어사이드 캐시의 성질).
---
## 5. 이것이 `persistent-user-sessions` 를 켜는 진짜 이유다
| | persistent (KC 26 기본) | volatile (KC 24 이전 방식) |
|---|---|---|
| 롤링 재시작 후 | **세션 유지** | **전원 로그아웃** |
| 배포 빈도 | 자유롭다 | 배포가 곧 사고다 |
| 대가 | DB 쓰기 (A-6 에서 본 지연) | 없음 |
**A-7 에서 volatile 로 바꿔 같은 실험을 반복하면 여기가 정반대가 될 것이다.**
그 비교가 이 실험의 짝이다.
---
---
## 개념
### StatefulSet 롤링 재시작의 무중단 조건
```
한 번에 하나씩 내린다 + readiness 로 전환 시점을 맞춘다
└─ 항상 최소 하나는 Ready 다
```
**두 가지가 다 있어야 성립한다.** replica 1 이면 반드시 끊기고,
readiness 프로브가 없으면 아직 기동 중인 파드로 트래픽이 간다.
### 룩어사이드 캐시가 재시작을 견디는 이유
| | 재시작 후 |
|---|---|
| 캐시 (프로세스 메모리) | **사라진다** |
| DB (진실의 원천) | 남는다 |
| 정확성 | **유지된다** — 첫 접근만 느려진다 |
A-0 에서 세운 모델이 여기서 그대로 확인된다.
---
---
## 증거 파일
**증거 수집 시각: 2026-09-04 13:19 13:20 KST** (파일 mtime 기준. 문서 상단의 시각 표기는 작성 시점이라 다를 수 있다.)
| 파일 | 종류 |
|---|---|
| [`01-restart-availability.txt`](evidence/a8-rolling-restart/01-restart-availability.txt) | 터미널 원문 |
| [`02-session-survival.txt`](evidence/a8-rolling-restart/02-session-survival.txt) | 터미널 원문 |
| [`a8-cache-reset-cluster-reformed.png`](evidence/a8-rolling-restart/a8-cache-reset-cluster-reformed.png) | 스크린샷 |
파일별 상세는 [`evidence/a8-rolling-restart/README.md`](evidence/a8-rolling-restart/README.md).
## 6. 재현 절차 (명령어)
```bash
# 상주 탐침 (StatefulSet 밖에 있어야 한다)
kubectl -n keycloak-lab run a8-probe --image=curlimages/curl:8.11.1 \
--restart=Never --command -- sleep 3600
kubectl -n keycloak-lab wait --for=condition=Ready pod/a8-probe --timeout=120s
# 로그인하고 토큰 보관
kubectl -n keycloak-lab exec a8-probe -- sh -c \
'curl -s -X POST http://<pod>:8080/realms/master/protocol/openid-connect/token \
-d grant_type=password -d client_id=admin-cli -d username=admin -d password=<pw> > /tmp/tok'
# 재시작 + 가용성 감시
kubectl -n keycloak-lab rollout restart statefulset/keycloak
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=420s
# 세션 생존 확인
kubectl -n keycloak-lab exec a8-probe -- sh -c \
'curl -s -o /dev/null -w "%{http_code}\n" -X POST http://<pod>:8080/realms/master/protocol/openid-connect/token \
-d grant_type=refresh_token -d client_id=admin-cli -d refresh_token=$(cat /tmp/rt)'
# DB 대조
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "select user_session_id, created_on, last_session_refresh from offline_user_session where user_session_id='<sid>'"
```
---
## 7. 다음 실험에 남기는 것
| 실험 | 이 실험이 준 것 |
|---|---|
| **A-7** volatile 비교 | **이 실험을 그대로 반복하면 정반대 결과가 나와야 한다** |
| **D-2** 버전 업그레이드 | 롤링 재시작이 안전하다는 것이 업그레이드의 전제 |
| 구성 | 무중단은 **replica ≥ 2 + readiness** 의 조합이다 |