docs: B-5 — the pod stays Ready while every request hangs
Stopping Redis returns HTTP 000 rather than an error because the client waits on reconnect, and the pod keeps serving traffic because the redis health indicator is not in the readiness group even though /actuator/health returns 503. That is the mirror image of A-2, where Keycloak put its database check in readiness and the pods left the Service. Turning on AOF with config set created the appendonlydir and still lost everything on pod deletion, because /data was the container filesystem; adding a PVC makes the same setting work. Volume first, persistence setting second. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
7dc0a3e5da
commit
f45a2a2aaa
@@ -0,0 +1,9 @@
|
||||
=== 기준선 ===
|
||||
Redis 키: 1
|
||||
PostgreSQL 토큰: 1 행
|
||||
Redis 영속화 설정:
|
||||
save = save
|
||||
appendonly no
|
||||
|
||||
=== 외부 진입점 정상 확인 ===
|
||||
https://app1.hyeonworks.com/ HTTP 200
|
||||
@@ -0,0 +1,25 @@
|
||||
=== ① Redis 정지 ===
|
||||
정지: 14:26:30
|
||||
deployment.apps/redis scaled
|
||||
삭제 완료
|
||||
|
||||
=== 로그인한 사용자의 다음 요청은 어떻게 되는가 ===
|
||||
/ HTTP 200
|
||||
/bff/token-boundary HTTP 000
|
||||
/actuator/health HTTP 503
|
||||
--- token-boundary 응답 본문 ---
|
||||
|
||||
|
||||
=== 파드 상태 — readiness 가 Redis 를 보는가 ===
|
||||
bff-555df79c97-6j86w 1/1 Running 0 17m
|
||||
bff-555df79c97-vgg6g 1/1 Running 0 16m
|
||||
|
||||
=== health 상세 ===
|
||||
|
||||
|
||||
=== BFF 로그 ===
|
||||
at java.base/sun.nio.ch.Net.pollConnect(Native Method) ~[na:na]
|
||||
at java.base/sun.nio.ch.Net.pollConnectNow(Unknown Source) ~[na:na]
|
||||
at java.base/sun.nio.ch.SocketChannelImpl.finishConnect(Unknown Source) ~[na:na]
|
||||
at io.netty.channel.socket.nio.NioSocketChannel.doFinishConnect(NioSocketChannel.java:336) ~[netty-transport-4.1.135.Final.jar!/:4.1.135.Final]
|
||||
at io.netty.channel.nio.AbstractNioChannel$AbstractNioUnsafe.finishConnect(AbstractNioChannel.java:339) ~[netty-transport-4.1.135.Final.jar!/:4.1.135.Final]
|
||||
@@ -0,0 +1,13 @@
|
||||
=== health 그룹별 응답 — 왜 파드는 Ready 인가 ===
|
||||
/actuator/health HTTP server
|
||||
/actuator/health/readiness HTTP 200
|
||||
/actuator/health/liveness HTTP 200
|
||||
|
||||
=== /actuator/health 본문 (Redis 항목이 있는가) ===
|
||||
|
||||
|
||||
=== /actuator/health/readiness 본문 ===
|
||||
{"status":"UP"}
|
||||
|
||||
=== Service 엔드포인트 — 트래픽을 계속 받는가 ===
|
||||
ready: [10.42.0.52 10.42.1.124]
|
||||
@@ -0,0 +1,39 @@
|
||||
=== 복구 ===
|
||||
deployment.apps/redis scaled
|
||||
deployment "redis" successfully rolled out
|
||||
/actuator/health HTTP 200
|
||||
/bff/token-boundary HTTP 302
|
||||
BFF 재시작 필요했나: 0,0 회 재시작
|
||||
|
||||
=== ② 영속화 — 지금 설정으로 재시작하면 무엇이 남는가 ===
|
||||
키 심음: before-restart
|
||||
dbsize: 4
|
||||
|
||||
--- AOF 를 켜고 다시 심는다 (영속화가 켜져 있으면 살아남는가) ---
|
||||
appendonly yes
|
||||
total 12
|
||||
drwxr-xr-x 3 redis redis 4096 Sep 4 05:26 .
|
||||
drwxr-xr-x 1 root root 4096 Sep 4 05:26 ..
|
||||
drwx------ 2 redis redis 4096 Sep 4 05:26 appendonlydir
|
||||
|
||||
--- 파드를 지운다 ---
|
||||
deployment "redis" successfully rolled out
|
||||
재기동 후:
|
||||
dbsize: 0
|
||||
b5:probe
|
||||
b5:aof
|
||||
appendonly no
|
||||
persistentvolumeclaim/redis-data created
|
||||
deployment.apps/redis configured
|
||||
deployment "redis" successfully rolled out
|
||||
|
||||
=== 영속 볼륨 위에서 다시 시험 ===
|
||||
appendonly yes
|
||||
키 심음: written-on-pvc
|
||||
sed: -e expression #1, char 8: unknown option to 's'
|
||||
|
||||
--- 파드를 지운다 ---
|
||||
deployment "redis" successfully rolled out
|
||||
재기동 후:
|
||||
dbsize: 1
|
||||
b5:pvc written-on-pvc
|
||||
@@ -0,0 +1,17 @@
|
||||
# B-5 — Redis 상실과 영속화 증거
|
||||
|
||||
2026-09-04 15:35–15:50 KST
|
||||
해설: [`docs/experiment-b5-redis-loss-persistence.md`](../../experiment-b5-redis-loss-persistence.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-baseline.txt` | 정지 전 — Redis 1키, PostgreSQL 1행, `save`/`appendonly no`, 외부 200 |
|
||||
| `02-redis-down.txt` | 정지 후 — `/bff/token-boundary` **`HTTP 000`(멈춤)**, `/actuator/health` 503, **파드는 1/1 Ready 유지**, Lettuce 재연결 스택 |
|
||||
| `03-health-groups.txt` | **핵심** — `/actuator/health` 503 인데 `/actuator/health/readiness` 는 `{"status":"UP"}`. Service 엔드포인트에 두 파드 모두 남아 있다 |
|
||||
| `04-persistence.txt` | 복구는 자동(재시작 0회) · **AOF 를 켰는데 파드 삭제 후 `dbsize 0`** · PVC 를 붙인 뒤 `written-on-pvc` **생존** |
|
||||
|
||||
## 핵심 세 줄
|
||||
|
||||
1. **파드가 Ready 를 유지한 채 계속 실패한다.** `redis` 헬스 지표가 readiness 그룹에 없기 때문이며, A-2 에서 Keycloak 이 NotReady 가 된 것과 정반대다.
|
||||
2. **오류가 아니라 멈춤이다.** `HTTP 000` — 빠른 실패가 안 되어 있어 사용자는 멈춘 화면을 본다.
|
||||
3. **볼륨 없이 AOF 만 켜는 것은 장식이다.** `appendonlydir` 까지 만들어지지만 컨테이너와 함께 사라진다.
|
||||
Reference in New Issue
Block a user