Compare commits

..
Author SHA1 Message Date
DongHyeonkaandClaude Opus 5 99b689e715 docs: A-3 — four logins returned tokens for sessions the crash erased
Keycloak commits the login INSERT with synchronous_commit off, so a crash loses whole sessions and not just refresh timestamps. Measured 4 of 153 lost, matching the default wal_writer_delay window.

Two injections failed silently first: --grace-period=0 --force lets the container runtime send SIGTERM so PostgreSQL flushes and shuts down cleanly, and SIGKILL to PID 1 from inside its own namespace is ignored by the kernel. Killing a backend makes the postmaster reinitialize, which is a real crash recovery.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 12:05:08 +09:00
DongHyeonkaandClaude Opus 5 4177fb6a48 docs: A-2 — losing the database takes every node down while up stays 1
Both pods go NotReady, the Service endpoint list empties and the front door returns 503, so adding Keycloak replicas buys nothing against database loss. The node holding the session in cache fails too, because a refresh writes LAST_SESSION_REFRESH. Recovery was automatic in about fifteen seconds with no restart, which is what readiness rather than liveness buys.

The observability finding matters as much: up stayed at 1 through a total outage, so alerting on it would have caught nothing. kube-state-metrics is missing and pod readiness is therefore not recorded as a metric.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 11:57:55 +09:00
DongHyeonkaandClaude Opus 5 2a98ef1090 docs: A-1 — sessions survive a JGroups partition but logout invalidation does not
Cutting TCP 7800 leaves cross-node refresh working (200), confirming sessions travel through PostgreSQL rather than the cluster transport. Logout is the opposite: the database row is deleted but the other node answers from its stale local cache, so the A-0 conclusion that invalidation rides the database is corrected here.

Two things the plan did not anticipate: a NetworkPolicy cannot sever an established connection because conntrack accepts it before policy evaluation, and Keycloak reports the partition through its readiness probe so the split node removes itself from the Service.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 11:52:49 +09:00
DongHyeonkaandClaude Opus 5 9c3cde457e docs: close nine gaps found by re-reading every open question in full
Cross-checked each question's 남은 미지수, 다음 검증 and 제약 against the plan item by item. Adds B-0 (autoconfiguration actually chosen), B-6 (encryption key rotation) and B-7 (oauth2-proxy cookie secret rotation) as new experiments, plus lock-holder death, rotation-disabled comparison, partial-logout recovery, store latency and the Q4 design checklist. Restores the Redis persistence comparison and records the correct index URL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 11:16:59 +09:00
DongHyeonkaandClaude Opus 5 f1e8c35805 docs: plan all remaining experiments with architecture, injection points and predictions
Twenty experiments across four layers, each with a topology diagram marking where the fault goes in, the metrics to watch, a falsifiable prediction written before the run, and a pass/fail rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 10:50:23 +09:00
DongHyeonkaandClaude Opus 5 1de6108157 docs: add the prerequisite knowledge this lab assumes
Builds up from HTTP statelessness to why session storage location determines the operational response, so the measurements have context to land in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 10:24:21 +09:00
DongHyeonkaandClaude Opus 5 22d873eb4f docs: capture the SQL the other node actually runs, and correct the replication claim
PostgreSQL statement logging shows keycloak-1 reading and updating the session created on keycloak-0. The same transaction reveals optimistic locking via VERSION, SKIP LOCKED, and synchronous_commit turned off. Fixes the earlier concept note that credited Infinispan with cross-node propagation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 10:14:45 +09:00
DongHyeonkaandClaude Opus 5 e5ebaeb623 docs: prove sessions are shared by PostgreSQL, not Infinispan replication
Experiment 0 with three probes: cross-node refresh/logout, cache counter deltas around a single login, and cache entry ownership. Each node caches only what it handled; cache totals sum exactly to the database count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 09:58:10 +09:00
DongHyeonkaandClaude Opus 5 006da7d490 docs: record observability setup and add concept layers 10-13
Adds Kubernetes resources (StatefulSet, PVC, Secret, RBAC, placement), Keycloak clustering internals (Infinispan, JGroups), Prometheus concepts and virtualization operations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 09:21:45 +09:00
DongHyeonkaandClaude Opus 5 d5cc2b55a9 fix: grant nodes/proxy so kubelet metrics can be scraped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 09:14:51 +09:00
DongHyeonkaandClaude Opus 5 0bb0e0ac49 feat: add Prometheus, node-exporter and Grafana for fault-injection observability
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 09:13:41 +09:00
DongHyeonkaandClaude Opus 5 33878e8880 docs: record multi-node cluster setup, rationale and formation evidence
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 17:26:45 +09:00
DongHyeonkaandClaude Opus 5 6dce35ec83 feat: deploy Keycloak multi-node cluster with PostgreSQL
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 17:17:25 +09:00
87 changed files with 7217 additions and 40 deletions
@@ -0,0 +1,22 @@
[ 1289ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/login:0
[ 1417ms] [VERBOSE] [DOM] Input elements should have autocomplete attributes (suggested: "username"): (More info: https://goo.gl/9p2vKq) %o @ https://app2.hyeonworks.com/login:0
[ 9024ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 9130ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/:0
[ 9456ms] [WARNING] Deprecation warning: value provided is not in a recognized RFC2822 or ISO format. moment construction falls back to js Date(), which is not reliable across all browsers and versions. Non RFC2822/ISO date formats are discouraged. Please refer to http://momentjs.com/guides/#/warnings/js-date/ for more info.
Arguments:
[0] _isAMomentObject: true, _isUTC: false, _useUTC: false, _l: undefined, _i: Thu, 27 Aug 2026 13:03:49, _f: undefined, _strict: undefined, _locale: [object Object]
Error
at a.createFromInputFallback (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:3)
at an (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:25647)
at un (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29355)
at aa (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29221)
at on (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:28938)
at sa (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29715)
at A (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29748)
at a (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:621:89)
at f (https://app2.hyeonworks.com/public/build/3719.c065b2e146c4c8347d51.js:1:4635)
at u (https://app2.hyeonworks.com/public/build/322.177b4bb01c5d74f9b28f.js:2473:47448) @ https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:620
[ 9605ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 10620ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 11527ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 13875ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
@@ -0,0 +1,7 @@
[ 144ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 153ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore:0
[ 1077ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 2101ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 2922ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 7323ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 9370ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
@@ -0,0 +1,8 @@
[ 271ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&orgId=1&panes=%7B%22a%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22vendor_statistics_approximate_entries_unique%7Bcache%3D%5C%22sessions%5C%22%7D%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22%7B%7Bpod%7D%7D%20on%20%7B%7Bnode%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-15m%22%2C%22to%22%3A%22now%22%7D%7D%7D:0
[ 346ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 1512ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 2433ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 6941ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 13188ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 21578ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 25998ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
@@ -0,0 +1,7 @@
[ 766ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step1.html:0
[ 781ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step2.html:0
[ 17929ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step1.html:0
[ 17949ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step2.html:0
[ 17981ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/login-status-iframe.html:0
[ 18447ms] [WARNING] For accessibility reasons an aria-label should be specified on nav groups if a title isn't @ https://auth.hyeonworks.com/resources/9v5yc/admin/keycloak.v2/assets/main-BbID33M6.js:7
[ 18462ms] [WARNING] For accessibility reasons an aria-label should be specified on nav groups if a title isn't @ https://auth.hyeonworks.com/resources/9v5yc/admin/keycloak.v2/assets/main-BbID33M6.js:7
@@ -0,0 +1,2 @@
[ 75ms] [ERROR] Failed to load resource: the server responded with a status of 404 (Not Found) @ https://hyeonworks.com/questions:0
[ 100ms] [ERROR] Failed to load resource: the server responded with a status of 404 (Not Found) @ https://hyeonworks.com/favicon.ico:0
@@ -1 +1 @@
[ 295ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0 [ 149ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -1 +1 @@
[ 125ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0 [ 103ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -1 +1 @@
[ 91ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0 [ 132ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -1 +1 @@
[ 88ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0 [ 239ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -0,0 +1 @@
[ 105ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -0,0 +1 @@
[ 108ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
@@ -0,0 +1,41 @@
[ 344ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 869ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&panes=%7B%22cf1%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22vendor_cluster_size%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22%7B%7Bpod%7D%7D+on+%7B%7Bnode%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-45m%22%2C%22to%22%3A%22now%22%7D%7D%7D&orgId=1:0
[ 1014ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 2039ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 4085ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 7289ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 15556ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 26818ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 32659ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 35930ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 53849ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 61722ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 79450ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 84571ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 87233ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 92665ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 113244ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 119181ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 137390ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 157071ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 166091ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 182011ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 188613ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 206228ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 218561ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 229607ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 235818ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 237158ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 252095ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 261118ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 273809ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 280227ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 292650ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 306477ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 309957ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 311984ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 325683ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 344586ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 357576ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 359294ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 364740ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
@@ -0,0 +1,55 @@
[ 615ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&panes=%7B%22te0%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22up%7Bjob%3D%5C%22keycloak%5C%22%7D%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22up+%E2%80%94+%7B%7Bpod%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-20m%22%2C%22to%22%3A%22now%22%7D%7D%7D&orgId=1:0
[ 929ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 2566ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 5541ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 7990ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 12402ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 17640ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 20285ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 28277ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 34341ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 34985ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 50081ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 65649ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 75046ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 75890ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 96067ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 113467ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 121972ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 139681ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 150971ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 161211ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 180744ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 191401ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 210668ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 229082ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 239345ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 252119ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 266045ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 276705ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 289911ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 296875ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 315099ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 325431ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 328418ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 343774ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 347054ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 352405ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 361291ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 370639ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 374286ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 377413ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 394362ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 405423ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 425476ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 434299ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 441775ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 450786ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 456829ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 468812ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 478913ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 491241ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 495228ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 502295ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 511625ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
[ 524619ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
@@ -1,25 +0,0 @@
- generic [ref=f7e3]:
- link "본문으로 건너뛰기" [ref=f7e4] [cursor=pointer]:
- /url: "#main-content"
- banner [ref=f7e5]:
- generic [ref=f7e6]:
- link "TechLog 홈" [ref=f7e8] [cursor=pointer]:
- /url: /
- text: TechLog
- generic [ref=f7e9]:
- button "TechLog 검색 열기" [ref=f7e11] [cursor=pointer]: 검색
- group [ref=f7e12]:
- generic "메뉴" [ref=f7e13] [cursor=pointer]
- generic [ref=f7e14]:
- paragraph [ref=f7e15]: 화면을 준비하고 있습니다.
- generic [ref=f7e16]: TechLog 로딩 중
- contentinfo [ref=f7e17]:
- generic [ref=f7e18]:
- generic [ref=f7e19]:
- paragraph [ref=f7e20]: 동현
- paragraph [ref=f7e21]: 문제를 재현하고 검증해 실제 운영에 적용할 수 있는 형태로 정리합니다.
- generic [ref=f7e22]:
- link "프로필" [ref=f7e23] [cursor=pointer]:
- /url: /profile
- link "변경 기록" [ref=f7e24] [cursor=pointer]:
- /url: /releases
@@ -0,0 +1,2 @@
- main [ref=e7]:
- status "Loading" [ref=e10]
@@ -0,0 +1 @@
- main [ref=f3e7]
@@ -0,0 +1,2 @@
- main [ref=f6e7]:
- status "Loading" [ref=f6e10]
@@ -0,0 +1,104 @@
- generic [ref=f9e4]:
- link "Skip to main content" [ref=f9e5] [cursor=pointer]:
- /url: "#pageContent"
- banner [ref=f9e7]:
- generic [ref=f9e8]:
- link [ref=f9e10] [cursor=pointer]:
- /url: /
- img "Grafana" [ref=f9e11]
- generic [ref=f9e14]:
- button "Search or jump to..." [ref=f9e18] [cursor=pointer]
- generic [ref=f9e19]: ctrl+k
- generic [ref=f9e23]:
- button "New" [ref=f9e24] [cursor=pointer]
- button "Help" [ref=f9e30] [cursor=pointer]
- button "News" [ref=f9e33] [cursor=pointer]
- button "Profile" [ref=f9e36] [cursor=pointer]:
- img "User avatar" [ref=f9e37]
- generic [ref=f9e38]:
- button "Open menu" [ref=f9e40] [cursor=pointer]
- navigation "Breadcrumbs" [ref=f9e43]:
- list [ref=f9e44]:
- listitem [ref=f9e45]:
- link "Home" [ref=f9e46] [cursor=pointer]:
- /url: /
- listitem [ref=f9e50]:
- link "Explore" [ref=f9e51] [cursor=pointer]:
- /url: /explore
- listitem [ref=f9e55]:
- generic "Prometheus" [ref=f9e56]
- generic [ref=f9e57]:
- button "Show more items" [ref=f9e60] [cursor=pointer]
- button "Toggle top search bar" [ref=f9e64] [cursor=pointer]
- main [ref=f9e70]:
- generic [ref=f9e72]:
- heading "Explore" [level=1] [ref=f9e73]
- generic [ref=f9e78]:
- navigation "Explore toolbar" [ref=f9e80]:
- navigation "Search links" [ref=f9e82]:
- generic [ref=f9e83]:
- button "Content outline" [expanded] [ref=f9e85] [cursor=pointer]:
- generic [ref=f9e88]: Outline
- generic [ref=f9e93] [cursor=pointer]:
- img "Prometheus logo" [ref=f9e95]
- textbox "Select a data source" [ref=f9e96]:
- /placeholder: ""
- button "Show more items" [ref=f9e102] [cursor=pointer]
- generic [ref=f9e106]:
- generic [ref=f9e110]:
- button "Collapse outline" [expanded] [ref=f9e112] [cursor=pointer]:
- img "arrow-from-right" [ref=f9e113]
- button "Queries" [ref=f9e116] [cursor=pointer]:
- img "arrow" [ref=f9e117]
- generic [ref=f9e124]:
- generic [ref=f9e126]:
- generic "Query editor row" [ref=f9e129]:
- generic [ref=f9e130]:
- generic [ref=f9e132]:
- generic [ref=f9e133]:
- button "Collapse query row" [expanded] [ref=f9e134] [cursor=pointer]
- generic [ref=f9e137]:
- button "Query editor row title A" [ref=f9e138] [cursor=pointer]:
- generic [ref=f9e139]: A
- emphasis [ref=f9e140]: (Prometheus)
- generic [ref=f9e141]:
- button "Show data source help" [ref=f9e143] [cursor=pointer]
- button "Duplicate query" [ref=f9e147] [cursor=pointer]
- button "Hide response" [ref=f9e151] [cursor=pointer]
- button "Remove query" [ref=f9e155] [cursor=pointer]
- button "Drag and drop to reorder" [ref=f9e158]:
- img "Drag and drop to reorder" [ref=f9e159]
- generic [ref=f9e162]:
- generic [ref=f9e163]:
- button "Kick start your query" [ref=f9e164] [cursor=pointer]
- generic [ref=f9e167]:
- generic [ref=f9e168] [cursor=pointer]: Explain
- generic [ref=f9e169]:
- checkbox "Explain Toggle switch" [ref=f9e170]
- generic "Toggle switch" [ref=f9e171] [cursor=pointer]
- radiogroup [ref=f9e176]:
- generic [ref=f9e177]:
- radio "Builder" [ref=f9e178] [cursor=pointer]
- generic [ref=f9e179] [cursor=pointer]: Builder
- generic [ref=f9e180]:
- radio "Code" [checked] [ref=f9e181] [cursor=pointer]
- generic [ref=f9e182] [cursor=pointer]: Code
- generic [ref=f9e184]:
- generic [ref=f9e186]:
- button "Loading metrics..." [disabled] [ref=f9e187] [cursor=pointer]
- generic [ref=f9e190]: Loading editor
- 'button "Options Legend: {{pod}} on {{node}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=f9e198] [cursor=pointer]':
- generic [ref=f9e202]:
- heading "Options" [level=6] [ref=f9e203]
- generic [ref=f9e204]:
- generic [ref=f9e205]: "Legend: {{pod}} on {{node}}"
- generic [ref=f9e206]: "Format: Time series"
- generic [ref=f9e207]: "Step: auto"
- generic [ref=f9e208]: "Type: Range"
- generic [ref=f9e209]: "Exemplars: false"
- generic [ref=f9e210]:
- button "Add query" [ref=f9e211] [cursor=pointer]
- button "Query history" [ref=f9e215] [cursor=pointer]
- button "Query inspector" [ref=f9e219] [cursor=pointer]
- generic:
- main
@@ -0,0 +1,4 @@
- main [ref=f12e3]:
- generic [ref=f12e4]:
- progressbar "Contents" [ref=f12e5]
- paragraph [ref=f12e8]: Loading the Administration Console
@@ -0,0 +1,4 @@
- generic [active] [ref=f15e1]:
- progressbar "Loading" [ref=f15e4]
- generic:
- list
@@ -0,0 +1 @@
- generic [active] [ref=f18e1]: Not Found
@@ -0,0 +1,151 @@
- generic [active] [ref=e1]:
- generic [ref=e4]:
- link "Skip to main content" [ref=e5] [cursor=pointer]:
- /url: "#pageContent"
- banner [ref=e7]:
- generic [ref=e8]:
- link [ref=e10] [cursor=pointer]:
- /url: /
- img "Grafana" [ref=e11]
- generic [ref=e14]:
- button "Search or jump to..." [ref=e18] [cursor=pointer]
- generic [ref=e19]: ctrl+k
- generic [ref=e23]:
- button "New" [ref=e24] [cursor=pointer]
- button "Help" [ref=e30] [cursor=pointer]
- button "News" [ref=e33] [cursor=pointer]
- button "Profile" [ref=e36] [cursor=pointer]:
- img "User avatar" [ref=e37]
- generic [ref=e38]:
- button "Open menu" [ref=e40] [cursor=pointer]
- navigation "Breadcrumbs" [ref=e43]:
- list [ref=e44]:
- listitem [ref=e45]:
- link "Home" [ref=e46] [cursor=pointer]:
- /url: /
- listitem [ref=e50]:
- link "Explore" [ref=e51] [cursor=pointer]:
- /url: /explore
- listitem [ref=e55]:
- generic "Prometheus" [ref=e56]
- generic [ref=e57]:
- generic [ref=e60]:
- button "Copy shortened URL" [ref=e61] [cursor=pointer]
- button "Open copy link options" [ref=e64] [cursor=pointer]
- button "Toggle top search bar" [ref=e68] [cursor=pointer]
- main [ref=e74]:
- generic [ref=e76]:
- heading "Explore" [level=1] [ref=e77]
- generic [ref=e82]:
- navigation "Explore toolbar" [ref=e84]:
- navigation "Search links" [ref=e86]:
- generic [ref=e87]:
- button "Content outline" [expanded] [ref=e89] [cursor=pointer]:
- generic [ref=e92]: Outline
- generic [ref=e97] [cursor=pointer]:
- img "Prometheus logo" [ref=e99]
- textbox "Select a data source" [ref=e100]:
- /placeholder: ""
- generic [ref=e104]:
- button "Split the pane" [ref=e106] [cursor=pointer]:
- generic [ref=e109]: Split
- button "Add" [ref=e111] [cursor=pointer]
- generic [ref=e116]:
- 'button "Time range selected: Last 45 minutes" [ref=e117] [cursor=pointer]'
- button "Zoom out time range" [ref=e122] [cursor=pointer]
- generic [ref=e126]:
- button "Run query" [ref=e127] [cursor=pointer]
- button "Auto refresh turned off. Choose refresh time interval" [ref=e131] [cursor=pointer]
- generic [ref=e135]:
- generic [ref=e139]:
- button "Collapse outline" [expanded] [ref=e141] [cursor=pointer]:
- img "arrow-from-right" [ref=e142]
- button "Queries" [ref=e145] [cursor=pointer]:
- img "arrow" [ref=e146]
- button "Graph" [ref=e150] [cursor=pointer]:
- img "graph-bar" [ref=e151]
- generic [ref=e158]:
- generic [ref=e160]:
- generic "Query editor row" [ref=e163]:
- generic [ref=e164]:
- generic [ref=e166]:
- generic [ref=e167]:
- button "Collapse query row" [expanded] [ref=e168] [cursor=pointer]
- generic [ref=e171]:
- button "Query editor row title A" [ref=e172] [cursor=pointer]:
- generic [ref=e173]: A
- emphasis [ref=e174]: (Prometheus)
- generic [ref=e175]:
- button "Show data source help" [ref=e177] [cursor=pointer]
- button "Duplicate query" [ref=e181] [cursor=pointer]
- button "Hide response" [ref=e185] [cursor=pointer]
- button "Remove query" [ref=e189] [cursor=pointer]
- button "Drag and drop to reorder" [ref=e192]:
- img "Drag and drop to reorder" [ref=e193]
- generic [ref=e196]:
- generic [ref=e197]:
- button "Kick start your query" [ref=e198] [cursor=pointer]
- generic [ref=e201]:
- generic [ref=e202] [cursor=pointer]: Explain
- generic [ref=e203]:
- checkbox "Explain Toggle switch" [ref=e204]
- generic "Toggle switch" [ref=e205] [cursor=pointer]
- radiogroup [ref=e210]:
- generic [ref=e211]:
- radio "Builder" [ref=e212] [cursor=pointer]
- generic [ref=e213] [cursor=pointer]: Builder
- generic [ref=e214]:
- radio "Code" [checked] [ref=e215] [cursor=pointer]
- generic [ref=e216] [cursor=pointer]: Code
- generic [ref=e218]:
- generic [ref=e220]:
- button "Metrics browser" [ref=e221] [cursor=pointer]
- code [ref=e228]:
- generic [ref=e229]:
- generic [ref=e234]: vendor_cluster_size
- textbox "Editor content;Press Alt+F1 for Accessibility Options." [ref=e239]: vendor_cluster_size
- 'button "Options Legend: {{pod}} on {{node}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=e245] [cursor=pointer]':
- generic [ref=e249]:
- heading "Options" [level=6] [ref=e250]
- generic [ref=e251]:
- generic [ref=e252]: "Legend: {{pod}} on {{node}}"
- generic [ref=e253]: "Format: Time series"
- generic [ref=e254]: "Step: auto"
- generic [ref=e255]: "Type: Range"
- generic [ref=e256]: "Exemplars: false"
- generic [ref=e257]:
- button "Add query" [ref=e258] [cursor=pointer]
- button "Query history" [ref=e262] [cursor=pointer]
- button "Query inspector" [ref=e266] [cursor=pointer]
- main [ref=e270]:
- region [ref=e272]:
- generic [ref=e273]:
- heading "Graph" [level=2] [ref=e275]
- radiogroup [ref=e278]:
- generic [ref=e279]:
- radio "Lines" [checked] [ref=e280] [cursor=pointer]
- generic [ref=e281] [cursor=pointer]: Lines
- generic [ref=e282]:
- radio "Bars" [ref=e283] [cursor=pointer]
- generic [ref=e284] [cursor=pointer]: Bars
- generic [ref=e285]:
- radio "Points" [ref=e286] [cursor=pointer]
- generic [ref=e287] [cursor=pointer]: Points
- generic [ref=e288]:
- radio "Stacked lines" [ref=e289] [cursor=pointer]
- generic [ref=e290] [cursor=pointer]: Stacked lines
- generic [ref=e291]:
- radio "Stacked bars" [ref=e292] [cursor=pointer]
- generic [ref=e293] [cursor=pointer]: Stacked bars
- list [ref=e302]:
- listitem [ref=e303]:
- button "keycloak-0 on kc-lab-2" [ref=e307] [cursor=pointer]
- listitem [ref=e308]:
- button "keycloak-0 on kc-lab-2" [ref=e312] [cursor=pointer]
- listitem [ref=e313]:
- button "keycloak-1 on kc-lab-1" [ref=e317] [cursor=pointer]
- generic [ref=e322]:
- alert
- alert
- complementary
- complementary
@@ -0,0 +1,145 @@
- generic [active] [ref=f3e1]:
- generic [ref=f3e4]:
- link "Skip to main content" [ref=f3e5] [cursor=pointer]:
- /url: "#pageContent"
- banner [ref=f3e7]:
- generic [ref=f3e8]:
- link [ref=f3e10] [cursor=pointer]:
- /url: /
- img "Grafana" [ref=f3e11]
- generic [ref=f3e14]:
- button "Search or jump to..." [ref=f3e18] [cursor=pointer]
- generic [ref=f3e19]: ctrl+k
- generic [ref=f3e23]:
- button "New" [ref=f3e24] [cursor=pointer]
- button "Help" [ref=f3e30] [cursor=pointer]
- button "News" [ref=f3e33] [cursor=pointer]
- button "Profile" [ref=f3e36] [cursor=pointer]:
- img "User avatar" [ref=f3e37]
- generic [ref=f3e38]:
- button "Open menu" [ref=f3e40] [cursor=pointer]
- navigation "Breadcrumbs" [ref=f3e43]:
- list [ref=f3e44]:
- listitem [ref=f3e45]:
- link "Home" [ref=f3e46] [cursor=pointer]:
- /url: /
- listitem [ref=f3e50]:
- link "Explore" [ref=f3e51] [cursor=pointer]:
- /url: /explore
- listitem [ref=f3e55]:
- generic "Prometheus" [ref=f3e56]
- generic [ref=f3e57]:
- generic [ref=f3e60]:
- button "Copy shortened URL" [ref=f3e61] [cursor=pointer]
- button "Open copy link options" [ref=f3e64] [cursor=pointer]
- button "Toggle top search bar" [ref=f3e68] [cursor=pointer]
- main [ref=f3e74]:
- generic [ref=f3e76]:
- heading "Explore" [level=1] [ref=f3e77]
- generic [ref=f3e82]:
- navigation "Explore toolbar" [ref=f3e84]:
- navigation "Search links" [ref=f3e86]:
- generic [ref=f3e87]:
- button "Content outline" [expanded] [ref=f3e89] [cursor=pointer]:
- generic [ref=f3e92]: Outline
- generic [ref=f3e97] [cursor=pointer]:
- img "Prometheus logo" [ref=f3e99]
- textbox "Select a data source" [ref=f3e100]:
- /placeholder: ""
- generic [ref=f3e104]:
- button "Split the pane" [ref=f3e106] [cursor=pointer]:
- generic [ref=f3e109]: Split
- button "Add" [ref=f3e111] [cursor=pointer]
- generic [ref=f3e116]:
- 'button "Time range selected: Last 20 minutes" [ref=f3e117] [cursor=pointer]'
- button "Zoom out time range" [ref=f3e122] [cursor=pointer]
- generic [ref=f3e126]:
- button "Run query" [ref=f3e127] [cursor=pointer]
- button "Auto refresh turned off. Choose refresh time interval" [ref=f3e131] [cursor=pointer]
- generic [ref=f3e135]:
- generic [ref=f3e139]:
- button "Collapse outline" [expanded] [ref=f3e141] [cursor=pointer]:
- img "arrow-from-right" [ref=f3e142]
- button "Queries" [ref=f3e145] [cursor=pointer]:
- img "arrow" [ref=f3e146]
- button "Graph" [ref=f3e150] [cursor=pointer]:
- img "graph-bar" [ref=f3e151]
- generic [ref=f3e158]:
- generic [ref=f3e160]:
- generic "Query editor row" [ref=f3e163]:
- generic [ref=f3e164]:
- generic [ref=f3e166]:
- generic [ref=f3e167]:
- button "Collapse query row" [expanded] [ref=f3e168] [cursor=pointer]
- generic [ref=f3e171]:
- button "Query editor row title A" [ref=f3e172] [cursor=pointer]:
- generic [ref=f3e173]: A
- emphasis [ref=f3e174]: (Prometheus)
- generic [ref=f3e175]:
- button "Show data source help" [ref=f3e177] [cursor=pointer]
- button "Duplicate query" [ref=f3e181] [cursor=pointer]
- button "Hide response" [ref=f3e185] [cursor=pointer]
- button "Remove query" [ref=f3e189] [cursor=pointer]
- button "Drag and drop to reorder" [ref=f3e192]:
- img "Drag and drop to reorder" [ref=f3e193]
- generic [ref=f3e196]:
- generic [ref=f3e197]:
- button "Kick start your query" [ref=f3e198] [cursor=pointer]
- generic [ref=f3e201]:
- generic [ref=f3e202] [cursor=pointer]: Explain
- generic [ref=f3e203]:
- checkbox "Explain Toggle switch" [ref=f3e204]
- generic "Toggle switch" [ref=f3e205] [cursor=pointer]
- radiogroup [ref=f3e210]:
- generic [ref=f3e211]:
- radio "Builder" [ref=f3e212] [cursor=pointer]
- generic [ref=f3e213] [cursor=pointer]: Builder
- generic [ref=f3e214]:
- radio "Code" [checked] [ref=f3e215] [cursor=pointer]
- generic [ref=f3e216] [cursor=pointer]: Code
- generic [ref=f3e218]:
- generic [ref=f3e220]:
- button "Loading metrics..." [disabled] [ref=f3e221] [cursor=pointer]
- code [ref=f3e228]:
- generic [ref=f3e229]:
- generic [ref=f3e234]: "up{job=\"keycloak\"}"
- textbox "Editor content;Press Alt+F1 for Accessibility Options." [ref=f3e239]: "up{job=\"keycloak\"}"
- 'button "Options Legend: up — {{pod}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=f3e245] [cursor=pointer]':
- generic [ref=f3e249]:
- heading "Options" [level=6] [ref=f3e250]
- generic [ref=f3e251]:
- generic [ref=f3e252]: "Legend: up — {{pod}}"
- generic [ref=f3e253]: "Format: Time series"
- generic [ref=f3e254]: "Step: auto"
- generic [ref=f3e255]: "Type: Range"
- generic [ref=f3e256]: "Exemplars: false"
- generic [ref=f3e257]:
- button "Add query" [ref=f3e258] [cursor=pointer]
- button "Query history" [ref=f3e262] [cursor=pointer]
- button "Query inspector" [ref=f3e266] [cursor=pointer]
- main [ref=f3e270]:
- region [ref=f3e272]:
- generic [ref=f3e273]:
- heading "Graph" [level=2] [ref=f3e275]
- radiogroup [ref=f3e278]:
- generic [ref=f3e279]:
- radio "Lines" [checked] [ref=f3e280] [cursor=pointer]
- generic [ref=f3e281] [cursor=pointer]: Lines
- generic [ref=f3e282]:
- radio "Bars" [ref=f3e283] [cursor=pointer]
- generic [ref=f3e284] [cursor=pointer]: Bars
- generic [ref=f3e285]:
- radio "Points" [ref=f3e286] [cursor=pointer]
- generic [ref=f3e287] [cursor=pointer]: Points
- generic [ref=f3e288]:
- radio "Stacked lines" [ref=f3e289] [cursor=pointer]
- generic [ref=f3e290] [cursor=pointer]: Stacked lines
- generic [ref=f3e291]:
- radio "Stacked bars" [ref=f3e292] [cursor=pointer]
- generic [ref=f3e293] [cursor=pointer]: Stacked bars
- generic [ref=f3e294]: Loading plugin panel...
- generic [ref=f3e299]:
- alert
- alert
- complementary
- complementary
@@ -0,0 +1,46 @@
# Experiment A-1 — cut the JGroups transport (TCP 7800) while leaving discovery alone.
#
# The point is to separate two things that are easy to conflate:
#
# discovery how the nodes FIND each other -> PostgreSQL JGROUPS_PING table
# transport how they actually TALK -> TCP 7800
#
# Blocking only the transport produces a state that cannot happen on a single
# node: both members stay registered in the database, so each believes the other
# exists, yet no message gets through.
#
# kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
# kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
#
# NetworkPolicy is an ALLOWLIST, not a firewall with deny rules. There is no way
# to write "deny 7800". The moment a pod is selected by a policy carrying
# policyTypes: [Ingress], every inbound port is denied unless a rule permits it.
# So 7800 is blocked by *omission*: 8080 and 9000 are listed, 7800 is not.
#
# That makes the two allow rules load-bearing — get them wrong and the experiment
# measures a dead Keycloak instead of a partitioned cluster:
#
# 8080 the HTTP endpoint. Traefik, the other pod's REST calls, and the probe
# traffic all arrive here.
# 9000 the management port: /health/started, /health/ready, /health/live and
# /metrics. Losing it means the kubelet fails the readiness probe and
# kills the pod — the cluster would break for the wrong reason.
#
# Both rules deliberately omit `from:`, which allows those ports from any source.
# Narrowing the source is not the subject here; the 2-hop experiment already
# established how to do that by label when it matters.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: a1-block-jgroups-transport
namespace: keycloak-lab
spec:
podSelector:
matchLabels:
app: keycloak
policyTypes: [Ingress]
ingress:
- ports:
- { port: 8080, protocol: TCP } # HTTP — must stay open
- { port: 9000, protocol: TCP } # health + metrics — must stay open
# 7800 is absent on purpose. That is the whole experiment.
+277
View File
@@ -0,0 +1,277 @@
# Keycloak multi-node cluster with PostgreSQL.
#
# Goal of this manifest: two Keycloak pods on two different nodes must discover
# each other and form one Infinispan cluster. Keycloak 26 discovers peers through
# the database (jdbc-ping) rather than multicast, writing to a JGROUPS_PING table,
# but the cluster traffic itself runs over TCP 7800 between the pods. Those are
# two separate mechanisms, which is why "registered in the DB but not clustered"
# is a real failure mode — and one that a single node cannot reproduce.
#
# kubectl apply -f deploy/lab/k8s/keycloak-cluster.yaml
# kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=600s
#
# Secrets are plain here. Proper secret handling is roadmap item 11; keeping it
# visible for now is deliberate so the gap is obvious rather than forgotten.
apiVersion: v1
kind: Namespace
metadata:
name: keycloak-lab
---
apiVersion: v1
kind: Secret
metadata:
name: keycloak-lab-secrets
namespace: keycloak-lab
type: Opaque
stringData:
POSTGRES_PASSWORD: lab-postgres-change-me
KC_BOOTSTRAP_ADMIN_PASSWORD: lab-admin-change-me
---
# PostgreSQL. local-path binds the volume to whichever node the pod lands on, so
# the database is effectively pinned to one node. That is not a flaw here: it is
# what makes "the database node dies" a meaningful experiment later.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-data
namespace: keycloak-lab
spec:
accessModes: [ReadWriteOnce]
storageClassName: local-path
resources:
requests:
storage: 5Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: postgres
namespace: keycloak-lab
spec:
replicas: 1
strategy:
type: Recreate # RWO volume cannot be mounted by two pods at once
selector:
matchLabels:
app: postgres
template:
metadata:
labels:
app: postgres
spec:
containers:
- name: postgres
image: postgres:16-alpine
ports:
- containerPort: 5432
name: postgres
env:
- name: POSTGRES_DB
value: keycloak
- name: POSTGRES_USER
value: keycloak
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: keycloak-lab-secrets
key: POSTGRES_PASSWORD
# The image refuses to initialise into a non-empty mount, and
# local-path volumes are clean, but this keeps the data one level
# down so a lost+found or similar never blocks initdb.
- name: PGDATA
value: /var/lib/postgresql/data/pgdata
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
readinessProbe:
exec:
command: ["sh", "-c", "pg_isready -U keycloak -d keycloak"]
initialDelaySeconds: 10
periodSeconds: 5
resources:
requests:
memory: 192Mi
cpu: 50m
limits:
memory: 512Mi
volumes:
- name: data
persistentVolumeClaim:
claimName: postgres-data
---
apiVersion: v1
kind: Service
metadata:
name: postgres
namespace: keycloak-lab
spec:
selector:
app: postgres
ports:
- port: 5432
targetPort: postgres
---
# Keycloak. A StatefulSet rather than a Deployment so each pod keeps a stable
# name (keycloak-0, keycloak-1); cluster membership is far easier to read in
# logs and in the JGROUPS_PING table when the identities do not churn.
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: keycloak
namespace: keycloak-lab
spec:
serviceName: keycloak-headless
replicas: 2
podManagementPolicy: Parallel # both pods start together, so they race to
# register — which is the interesting case
selector:
matchLabels:
app: keycloak
template:
metadata:
labels:
app: keycloak
spec:
# One pod per node. Two pods on one node would share a kernel and make the
# 7800 blocking experiment meaningless.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: keycloak
containers:
- name: keycloak
image: quay.io/keycloak/keycloak:26.7.0
# "start", not "start-dev". Dev mode forces cache=local and there is
# no cluster to form at all.
args: ["start"]
ports:
- containerPort: 8080
name: http
- containerPort: 9000
name: management
- containerPort: 7800
name: jgroups
env:
- name: KC_DB
value: postgres
- name: KC_DB_URL
value: jdbc:postgresql://postgres:5432/keycloak
- name: KC_DB_USERNAME
value: keycloak
- name: KC_DB_PASSWORD
valueFrom:
secretKeyRef:
name: keycloak-lab-secrets
key: POSTGRES_PASSWORD
# Settings confirmed by the two-hop header measurement.
# KC_HOSTNAME carries the full external URL, which pins scheme and
# host for issuer and redirect URLs regardless of headers.
# KC_PROXY_HEADERS is the separate opt-in that lets the forwarded
# client address through — the same kind of switch as Spring's
# forward-headers-strategy. See docs/two-hop-proxy-header-contract.md.
- name: KC_HOSTNAME
value: https://auth.hyeonworks.com
- name: KC_HOSTNAME_STRICT
value: "true"
- name: KC_PROXY_HEADERS
value: xforwarded
- name: KC_HTTP_ENABLED
value: "true"
- name: KC_HEALTH_ENABLED
value: "true"
- name: KC_METRICS_ENABLED
value: "true"
# Without an explicit cap the JVM sizes its heap from the container
# limit and this lab has roughly 3.8GB of guest headroom in total.
- name: JAVA_OPTS_KC_HEAP
value: "-Xms256m -Xmx512m"
- name: KC_BOOTSTRAP_ADMIN_USERNAME
value: admin
- name: KC_BOOTSTRAP_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: keycloak-lab-secrets
key: KC_BOOTSTRAP_ADMIN_PASSWORD
# Keycloak serves health and metrics on the management port (9000),
# not on 8080, since version 25.
startupProbe:
httpGet:
path: /health/started
port: management
periodSeconds: 10
failureThreshold: 60 # first boot runs an implicit build
readinessProbe:
httpGet:
path: /health/ready
port: management
periodSeconds: 10
livenessProbe:
httpGet:
path: /health/live
port: management
periodSeconds: 30
resources:
requests:
memory: 640Mi
cpu: 100m
limits:
memory: 900Mi
---
# Headless service. Not required for jdbc-ping discovery, which goes through the
# database, but it gives each pod a stable DNS name for direct inspection.
apiVersion: v1
kind: Service
metadata:
name: keycloak-headless
namespace: keycloak-lab
spec:
clusterIP: None
selector:
app: keycloak
ports:
- port: 8080
targetPort: http
name: http
- port: 9000
targetPort: management
name: management
---
apiVersion: v1
kind: Service
metadata:
name: keycloak
namespace: keycloak-lab
spec:
selector:
app: keycloak
ports:
- port: 8080
targetPort: http
name: http
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: keycloak
namespace: keycloak-lab
spec:
ingressClassName: traefik
rules:
- host: auth.hyeonworks.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: keycloak
port:
number: 8080
+373
View File
@@ -0,0 +1,373 @@
# Prometheus + node-exporter + Grafana.
#
# Purpose: during a fault-injection experiment, know *which signal moved first*.
# Without a metrics store the only record is whatever scrolled past in a terminal,
# and "the cluster recovered in about a minute" is not a measurement.
#
# kubectl apply -f deploy/lab/k8s/observability.yaml
# kubectl -n observability rollout status deployment/prometheus --timeout=300s
#
# Placement decision — Prometheus and Grafana are pinned to the control-plane
# node (kc-lab-1). An observability stack must not share a failure domain with
# the thing it observes. With only two nodes that cannot be fully avoided, so the
# rule here is: the node that gets killed in experiments is the *agent*
# (kc-lab-2, holding keycloak-0 and postgres), and everything needed to watch
# that happen lives on the server node.
apiVersion: v1
kind: Namespace
metadata:
name: observability
---
# Prometheus discovers scrape targets by querying the Kubernetes API, so it
# needs read access to nodes, services, endpoints and pods. Without this the
# kubernetes_sd_configs below silently return no targets.
apiVersion: v1
kind: ServiceAccount
metadata:
name: prometheus
namespace: observability
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: prometheus
rules:
- apiGroups: [""]
# nodes/proxy is required in addition to nodes/metrics: the kubelet job
# reaches each node through the API server's proxy subresource
# (/api/v1/nodes/<name>/proxy/metrics). Without it every kubelet target
# fails with 403 Forbidden while the other jobs stay green — a partial
# failure that is easy to miss unless the target list is checked.
resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
verbs: [get, list, watch]
- nonResourceURLs: ["/metrics"]
verbs: [get]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: prometheus
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: prometheus
subjects:
- kind: ServiceAccount
name: prometheus
namespace: observability
---
apiVersion: v1
kind: ConfigMap
metadata:
name: prometheus-config
namespace: observability
data:
prometheus.yml: |
global:
# 15s is short for production but right here: a node loss should show up
# within a couple of samples, not a minute later.
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
# Prometheus scraping itself. Useful as a control: if this target is down,
# the problem is Prometheus, not the thing being measured.
- job_name: prometheus
static_configs:
- targets: ['localhost:9090']
# Keycloak. Metrics live on the management port 9000, not 8080 — the same
# split that the health probes use. KC_METRICS_ENABLED=true is already set
# on the StatefulSet.
#
# Discovery is by endpoints rather than a static list because pod IPs
# change on every restart; that was observed directly when the lab was
# power-cycled and every pod came back with a new address.
- job_name: keycloak
kubernetes_sd_configs:
- role: endpoints
namespaces:
names: [keycloak-lab]
relabel_configs:
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
action: keep
regex: keycloak-headless;management
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: node
# node-exporter, one per node via DaemonSet. This is what answers
# "did the machine die or did the process die".
- job_name: node-exporter
kubernetes_sd_configs:
- role: endpoints
namespaces:
names: [observability]
relabel_configs:
- source_labels: [__meta_kubernetes_service_name]
action: keep
regex: node-exporter
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: node
# The kubelet's own metrics, reached through the API server proxy so no
# extra port needs opening.
- job_name: kubelet
scheme: https
tls_config:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecure_skip_verify: true
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: node
relabel_configs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- target_label: __address__
replacement: kubernetes.default.svc:443
- source_labels: [__meta_kubernetes_node_name]
regex: (.+)
target_label: __metrics_path__
replacement: /api/v1/nodes/${1}/proxy/metrics
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: prometheus-data
namespace: observability
spec:
accessModes: [ReadWriteOnce]
storageClassName: local-path
resources:
requests:
storage: 5Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: prometheus
namespace: observability
spec:
replicas: 1
strategy:
type: Recreate # RWO volume; two pods cannot mount it at once
selector:
matchLabels:
app: prometheus
template:
metadata:
labels:
app: prometheus
spec:
serviceAccountName: prometheus
# See the placement note at the top of this file.
nodeSelector:
node-role.kubernetes.io/control-plane: "true"
securityContext:
fsGroup: 65534 # the image runs as nobody and must own the volume
containers:
- name: prometheus
image: prom/prometheus:v3.1.0
args:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.path=/prometheus
# 7 days is far more than an experiment needs and keeps the volume
# small enough that it never becomes the reason a node fills up.
- --storage.tsdb.retention.time=7d
- --web.enable-lifecycle
ports:
- containerPort: 9090
name: http
volumeMounts:
- name: config
mountPath: /etc/prometheus
- name: data
mountPath: /prometheus
readinessProbe:
httpGet: { path: /-/ready, port: http }
initialDelaySeconds: 10
livenessProbe:
httpGet: { path: /-/healthy, port: http }
initialDelaySeconds: 30
resources:
requests: { memory: 256Mi, cpu: 50m }
limits: { memory: 640Mi }
volumes:
- name: config
configMap:
name: prometheus-config
- name: data
persistentVolumeClaim:
claimName: prometheus-data
---
apiVersion: v1
kind: Service
metadata:
name: prometheus
namespace: observability
spec:
selector:
app: prometheus
ports:
- port: 9090
targetPort: http
---
# node-exporter. A DaemonSet so every node reports, including one that is about
# to be killed — the last samples before it goes silent are the interesting part.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: observability
spec:
selector:
matchLabels:
app: node-exporter
template:
metadata:
labels:
app: node-exporter
spec:
# Host namespaces: the point is to measure the machine, not the container.
hostNetwork: true
hostPID: true
tolerations:
- operator: Exists # must also run on tainted nodes
containers:
- name: node-exporter
image: prom/node-exporter:v1.8.2
args:
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --path.rootfs=/host/root
- --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
ports:
- containerPort: 9100
name: metrics
hostPort: 9100
volumeMounts:
- { name: proc, mountPath: /host/proc, readOnly: true }
- { name: sys, mountPath: /host/sys, readOnly: true }
- { name: rootfs, mountPath: /host/root, readOnly: true, mountPropagation: HostToContainer }
resources:
requests: { memory: 32Mi, cpu: 20m }
limits: { memory: 96Mi }
volumes:
- { name: proc, hostPath: { path: /proc } }
- { name: sys, hostPath: { path: /sys } }
- { name: rootfs, hostPath: { path: / } }
---
apiVersion: v1
kind: Service
metadata:
name: node-exporter
namespace: observability
spec:
clusterIP: None # headless: Prometheus wants each pod, not a VIP
selector:
app: node-exporter
ports:
- port: 9100
targetPort: metrics
name: metrics
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: grafana
namespace: observability
spec:
replicas: 1
selector:
matchLabels:
app: grafana
template:
metadata:
labels:
app: grafana
spec:
nodeSelector:
node-role.kubernetes.io/control-plane: "true"
containers:
- name: grafana
image: grafana/grafana:11.4.0
ports:
- containerPort: 3000
name: http
env:
- name: GF_SECURITY_ADMIN_USER
value: admin
- name: GF_SECURITY_ADMIN_PASSWORD
value: lab-grafana-change-me
# Grafana builds absolute URLs for redirects and asset paths. Behind
# the nginx -> Traefik chain it must be told the external address,
# for exactly the reason Keycloak needs KC_HOSTNAME. Without it,
# login redirects come back as http://<pod-ip>:3000.
- name: GF_SERVER_ROOT_URL
value: https://app2.hyeonworks.com
volumeMounts:
- name: datasources
mountPath: /etc/grafana/provisioning/datasources
readinessProbe:
httpGet: { path: /api/health, port: http }
initialDelaySeconds: 15
resources:
requests: { memory: 128Mi, cpu: 50m }
limits: { memory: 320Mi }
volumes:
- name: datasources
configMap:
name: grafana-datasources
---
# Provisioning the datasource as a file means Grafana comes up already wired to
# Prometheus. Clicking through the UI would leave the configuration only in
# Grafana's own database, which is emptyDir here and disappears on restart.
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-datasources
namespace: observability
data:
prometheus.yaml: |
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus.observability.svc:9090
isDefault: true
---
apiVersion: v1
kind: Service
metadata:
name: grafana
namespace: observability
spec:
selector:
app: grafana
ports:
- port: 3000
targetPort: http
---
# Grafana is published on app2.hyeonworks.com because that name is already in
# the wildcard-free certificate (auth / app1 / app2) and is otherwise unused.
# It moves when app2 is needed for the SSO experiment.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: grafana
namespace: observability
spec:
ingressClassName: traefik
rules:
- host: app2.hyeonworks.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: grafana
port:
number: 3000
+55
View File
@@ -0,0 +1,55 @@
#!/usr/bin/env bash
# Experiment 0c — where does a session entry actually live?
#
# Experiment 0b showed keycloak-1's session cache never moved when keycloak-0
# handled a login. That leaves two explanations:
#
# (a) a DISTRIBUTED cache with owners=1 — entries are spread across nodes by
# consistent hashing, and this one happened to land on keycloak-0;
# (b) a LOCAL cache — each node only ever caches what it handled itself.
#
# They are distinguished by driving logins at the OTHER node. Under (a) the
# entries would keep landing on both nodes regardless of who was asked. Under
# (b) the count rises only on the node that received the request.
set -uo pipefail
NS="${NS:-keycloak-lab}"
N="${N:-5}"
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
echo " keycloak-0 = $K0_IP ($(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.spec.nodeName}'))"
echo " keycloak-1 = $K1_IP ($(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.spec.nodeName}'))"
echo
kubectl -n "$NS" run kc-own --rm -i --restart=Never \
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
O=/tmp/o; : > \$O
ent() {
curl -s --retry 3 --max-time 20 http://\$1:9000/metrics \
| grep -E '^vendor_statistics_approximate_entries_unique.cache=.sessions' \
| awk '{print \$NF}'
}
login() { i=0; while [ \$i -lt $N ]; do
curl -s -o /dev/null -X POST http://\$1:8080/realms/master/protocol/openid-connect/token \
-d grant_type=password -d client_id=admin-cli \
-d username=admin -d 'password=$ADMIN_PW'
i=\$((i+1)); done; sleep 5; }
{
printf '%-32s %12s %12s\n' '단계' 'k0 entries' 'k1 entries'
printf '%-32s %12s %12s\n' '시작' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
login $K1_IP
printf '%-32s %12s %12s\n' 'keycloak-1 에 로그인 ${N}회' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
login $K0_IP
printf '%-32s %12s %12s\n' 'keycloak-0 에 로그인 ${N}회' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
} >> \$O
cat \$O
" 2>&1 | grep -v '^pod .* deleted$'
echo
echo "=== 대조: PostgreSQL 에는 몇 건인가 ==="
kubectl -n "$NS" exec deploy/postgres -- psql -U keycloak -d keycloak -tAc \
"select count(*) from offline_user_session where offline_flag='0'" 2>/dev/null | sed 's/^/ online 세션 /'
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# Experiment 0b — does the Infinispan cache itself replicate, or do both nodes
# merely agree because they read the same database?
#
# Experiment 0 proved the two nodes give the same answers. That alone does NOT
# prove Infinispan replicated anything: with persistent-user-sessions (the
# Keycloak 26 default) the session is written to PostgreSQL, so two nodes reading
# one database would agree even with the cache disabled entirely.
#
# This script separates the two by measuring the cache counters on BOTH nodes
# around a single login. If the write on keycloak-0 shows up as cache activity
# on keycloak-1, the replication is real and not a database artifact.
set -uo pipefail
NS="${NS:-keycloak-lab}"
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
echo
# 파드 출력을 스트리밍으로 받으면 조각이 유실된다. 실제로 첫 시도에서
# keycloak-1 의 스냅샷과 그 다음 마커가 통째로 사라져 델타가 0 으로 보였다.
# 파드 안에서 파일로 모았다가 마지막에 한 번만 내보낸다.
kubectl -n "$NS" run kc-delta --rm -i --restart=Never \
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
set -u
K0='http://$K0_IP'; K1='http://$K1_IP'
O=/tmp/o.txt; : > \$O
snap() {
curl -s --retry 3 --retry-connrefused --max-time 20 \$1:9000/metrics \
| grep -E '^vendor_(statistics_(stores|hits|misses|approximate_entries_unique)|rpc_manager_replication_count)\{cache=\"(sessions|clientSessions)\"' \
| sed 's/,cache_manager=\"keycloak\"//; s/,node=\"[^\"]*\"//' >> \$O
}
echo '###BEFORE_K0' >> \$O; snap \$K0
echo '###BEFORE_K1' >> \$O; snap \$K1
echo '###LOGIN' >> \$O
curl -s -o /dev/null -w 'http_code=%{http_code}\n' -X POST \
\"\$K0:8080/realms/master/protocol/openid-connect/token\" \
-d grant_type=password -d client_id=admin-cli \
-d username=admin -d 'password=$ADMIN_PW' >> \$O
sleep 5
echo '###AFTER_K0' >> \$O; snap \$K0
echo '###AFTER_K1' >> \$O; snap \$K1
echo '###END' >> \$O
cat \$O
" 2>&1 | grep -v '^pod .* deleted$' > /tmp/cache-delta.txt
python3 - /tmp/cache-delta.txt <<'PY'
import re, sys
raw = open(sys.argv[1]).read()
blocks, cur = {}, None
for line in raw.splitlines():
if line.startswith('###'):
cur = line[3:]; blocks[cur] = {}
elif cur and '{' in line:
m = re.match(r'(\S+?)\{cache="(\w+)"\}\s+(\S+)', line)
if m:
blocks[cur][(m.group(1), m.group(2))] = float(m.group(3))
print('=== 로그인은 keycloak-0 에만 보냈다 ===')
code = [l for l in raw.splitlines() if l.startswith('http_code=')]
print(' 로그인 응답: ' + (code[0] if code else '없음'))
for n in ('BEFORE_K0','BEFORE_K1','AFTER_K0','AFTER_K1'):
if not blocks.get(n):
print(f' !! {n} 스냅샷이 비었다 — 델타를 신뢰할 수 없다')
print()
hdr = f" {'계수기':<42} {'캐시':<15} {'전':>8} {'후':>8} {'증가':>7}"
for node in ('K0', 'K1'):
who = 'keycloak-0 (로그인을 받은 노드)' if node == 'K0' else 'keycloak-1 (아무 요청도 받지 않은 노드)'
print(f'=== {who} ===')
print(hdr)
b, a = blocks.get(f'BEFORE_{node}', {}), blocks.get(f'AFTER_{node}', {})
for k in sorted(set(b) | set(a)):
before, after = b.get(k[0:2], 0.0), a.get(k[0:2], 0.0)
d = after - before
mark = ' ←' if d else ''
name = k[0].replace('vendor_statistics_', '').replace('vendor_rpc_manager_', 'rpc.')
print(f" {name:<42} {k[1]:<15} {before:>8.0f} {after:>8.0f} {d:>+7.0f}{mark}")
print()
PY
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env bash
# Experiment 0d — capture the actual SQL that the OTHER node runs.
#
# Experiments 0b/0c showed that session entries never appear in keycloak-1's
# memory, yet keycloak-1 can use a session keycloak-0 created. The conclusion
# "keycloak-1 reads it from PostgreSQL" was an inference, not an observation.
#
# This script turns on statement logging in PostgreSQL for a few seconds, sends
# ONE refresh request to keycloak-1 for a session born on keycloak-0, and greps
# the database log for that session id. If the inference is right, the SQL is
# there, issued from keycloak-1's pod IP.
#
# It also checks whether serving that request makes keycloak-1 cache the session
# — which sharpens "each node caches what it handled" from "what it logged in"
# to "what it touched".
set -uo pipefail
NS="${NS:-keycloak-lab}"
PSQL="kubectl -n $NS exec deploy/postgres -- psql -U keycloak -d keycloak -tAc"
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
echo " keycloak-0 = $K0_IP (세션을 만드는 노드)"
echo " keycloak-1 = $K1_IP (읽기만 하는 노드)"
echo
# %h 를 넣어야 어느 파드가 보낸 질의인지 로그에서 구분된다.
echo "=== PostgreSQL 문장 로깅을 켠다 ==="
$PSQL "alter system set log_statement='all'" >/dev/null 2>&1
$PSQL "alter system set log_line_prefix='%m [%p] %h '" >/dev/null 2>&1
$PSQL "select pg_reload_conf()" >/dev/null 2>&1
echo " log_statement = $($PSQL 'show log_statement' 2>/dev/null)"
echo " log_line_prefix = $($PSQL 'show log_line_prefix' 2>/dev/null)"
echo
# 로그 커서를 잡아둔다. 이 줄 수 이후만 본다.
LOG_BEFORE=$(kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null | wc -l)
RESULT=$(kubectl -n "$NS" run kc-readpath --rm -i --restart=Never \
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
O=/tmp/o; : > \$O
TOKEN_EP='/realms/master/protocol/openid-connect/token'
jget() { sed -n \"s/.*\\\"\$1\\\":\\\"\\([^\\\"]*\\)\\\".*/\\1/p\"; }
ent() {
curl -s --retry 3 --max-time 20 http://\$1:9000/metrics \
| grep -E '^vendor_statistics_approximate_entries_unique.cache=.sessions' | awk '{print \$NF}'
}
# keycloak-0 에서 로그인한다
L=\$(curl -s -X POST \"http://$K0_IP:8080\$TOKEN_EP\" -d grant_type=password \
-d client_id=admin-cli -d username=admin -d 'password=$ADMIN_PW')
SID=\$(echo \"\$L\" | jget access_token | cut -d. -f2 | sed 's/\$/==/' | base64 -d 2>/dev/null | jget sid)
RT=\$(echo \"\$L\" | jget refresh_token)
echo \"SID=\$SID\" >> \$O
echo \"K1_ENTRIES_BEFORE=\$(ent $K1_IP)\" >> \$O
sleep 2
# 반대편 노드에 refresh 를 딱 한 번 보낸다
# 인용을 한 겹 더 쌓으면 curl 이 URL 을 통째로 못 읽는다. 실제로 000 이 나왔다.
CODE=\$(curl -s -o /dev/null -w '%{http_code}' -X POST \
\"http://$K1_IP:8080\$TOKEN_EP\" \
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\")
echo \"REFRESH_ON_K1=\$CODE\" >> \$O
sleep 3
echo \"K1_ENTRIES_AFTER=\$(ent $K1_IP)\" >> \$O
cat \$O
" 2>&1 | grep -v '^pod .* deleted$')
echo "=== 요청 ==="
echo "$RESULT" | sed 's/^/ /'
SID=$(echo "$RESULT" | sed -n 's/^SID=//p')
echo
echo "=== PostgreSQL 문장 로깅을 끈다 ==="
$PSQL "alter system reset log_statement" >/dev/null 2>&1
$PSQL "alter system reset log_line_prefix" >/dev/null 2>&1
$PSQL "select pg_reload_conf()" >/dev/null 2>&1
echo " log_statement = $($PSQL 'show log_statement' 2>/dev/null)"
echo
echo "=== keycloak-1 이 실제로 보낸 SQL 문장 ==="
echo " (파라미터가 \$1 로 묶여 있어, sid 는 바로 아래 DETAIL 줄에 있다)"
echo
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
| tail -n +$((LOG_BEFORE + 1)) \
| grep -F "$K1_IP" | grep -E "LOG: execute" \
| sed 's/.*execute [^:]*: //' | sed 's/^/ /' | head -12
echo
echo "=== 그 sid 를 언급한 SQL — 누가 보냈는가 ==="
echo " 찾는 sid: $SID"
echo
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
| tail -n +$((LOG_BEFORE + 1)) \
| grep -F "$SID" \
| sed -e "s/$K0_IP/[keycloak-0]/g" -e "s/$K1_IP/[keycloak-1]/g" \
| cut -c1-220 \
| head -20
echo
echo "=== 요약: 파드별 질의 건수 ==="
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
| tail -n +$((LOG_BEFORE + 1)) \
| grep -F "$SID" \
| grep -oE "^[0-9-]+ [0-9:.]+ [A-Z]+ \[[0-9]+\] [0-9.]+" \
| awk '{print $NF}' | sort | uniq -c \
| sed -e "s/$K0_IP/[keycloak-0]/" -e "s/$K1_IP/[keycloak-1]/" -e 's/^/ /'
+207
View File
@@ -0,0 +1,207 @@
#!/usr/bin/env bash
# Experiment 0 — is a session created on one Keycloak node usable on the other?
#
# Forming a cluster is not the same as sharing session state. The Infinispan log
# says "cluster view (2)", but that only proves the members found each other.
#
# Design notes, learned the hard way:
#
# * Every probe has a CONTROL. A result from the far node means nothing unless
# the same call against the issuing node is also measured. The first version
# of this script reported "403 on keycloak-1" as if it were a replication
# failure; the issuing node returned 403 too, and the cause was a missing
# openid scope. Measure both, always.
#
# * Sessions are tracked by SID, not by count. Both the test login and the
# admin API calls create sessions for the same user, so counts are noisy.
# A specific session id either appears in a node's answer or it does not.
#
# * The probe is the REFRESH TOKEN grant, not userinfo. userinfo only validates
# a signature and can succeed on a node that knows nothing about the session.
# Refreshing requires the node to find the session, check it is alive, and
# write back a new refresh time — it actually touches the session store.
#
# Talks to pod IPs directly: going through nginx/Traefik would hide which node
# handled each request, which is the entire question.
#
# ./deploy/lab/scripts/experiment-session-replication.sh
set -uo pipefail
NS="${NS:-keycloak-lab}"
OUT="${OUT:-/tmp/session-replication}"
mkdir -p "$OUT"
PSQL="kubectl -n $NS exec deploy/postgres -- psql -U keycloak -d keycloak -tAc"
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
echo
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
K0_NODE=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.spec.nodeName}')
K1_NODE=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.spec.nodeName}')
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
echo "=== 대상 ==="
printf ' keycloak-0 %-14s %s\n' "$K0_IP" "$K0_NODE"
printf ' keycloak-1 %-14s %s\n' "$K1_IP" "$K1_NODE"
echo
echo "=== [0] 실험 전 DB 세션 ==="
$PSQL "select offline_flag, count(*) from offline_user_session group by offline_flag" 2>/dev/null \
| sed 's/^/ offline_flag=/' || echo " (없음)"
echo
# 파드 하나 안에서 전 단계를 실행한다. 단계마다 파드를 새로 띄우면 토큰을
# 단계 사이로 넘길 수 없다.
kubectl -n "$NS" run kc-probe --rm -i --restart=Never \
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
set -u
K0='http://$K0_IP:8080'; K1='http://$K1_IP:8080'
TOKEN_EP='/realms/master/protocol/openid-connect/token'
jget() { sed -n \"s/.*\\\"\$1\\\":\\\"\\([^\\\"]*\\)\\\".*/\\1/p\"; }
# ── [1] keycloak-0 에서 로그인. 이 노드가 세션의 출생지다 ──────────────────
LOGIN=\$(curl -s -X POST \"\$K0\$TOKEN_EP\" \
-d grant_type=password -d client_id=admin-cli \
-d username=admin -d 'password=$ADMIN_PW')
echo '###STEP1_LOGIN'; echo \"\$LOGIN\"
AT=\$(echo \"\$LOGIN\" | jget access_token)
RT=\$(echo \"\$LOGIN\" | jget refresh_token)
# ── [2] 관리 API 조회용 토큰. 세션 오염을 피하려고 따로 하나만 더 만든다 ──
ADMTOK=\$(curl -s -X POST \"\$K0\$TOKEN_EP\" \
-d grant_type=password -d client_id=admin-cli \
-d username=admin -d 'password=$ADMIN_PW' | jget access_token)
CID=\$(curl -s -H \"Authorization: Bearer \$ADMTOK\" \
\"\$K0/admin/realms/master/clients?clientId=admin-cli\" | jget id | head -1)
# ── [3] 두 노드에 같은 질문을 한다: admin-cli 의 세션 목록 ────────────────
echo '###STEP3_SESSIONS_K0'
curl -s -H \"Authorization: Bearer \$ADMTOK\" \
\"\$K0/admin/realms/master/clients/\$CID/user-sessions?max=100\"
echo
echo '###STEP3_SESSIONS_K1'
curl -s -H \"Authorization: Bearer \$ADMTOK\" \
\"\$K1/admin/realms/master/clients/\$CID/user-sessions?max=100\"
echo
# ── [4] 대조군: keycloak-0 이 발급한 refresh token 을 keycloak-0 에 쓴다 ──
# 먼저 반대편에 써야 하므로 여기서는 쓰지 않고, 순서를 [5] 뒤로 미룬다.
# refresh token 은 회전(rotation)되므로 한 번 쓰면 옛 것이 무효가 된다.
# 따라서 '반대편 먼저'가 유일하게 의미 있는 순서다.
# ── [5] 시험군: keycloak-0 이 발급한 refresh token 을 keycloak-1 에 쓴다 ──
echo '###STEP5_REFRESH_ON_K1'
curl -s -w '\nhttp_code=%{http_code}\n' -X POST \"\$K1\$TOKEN_EP\" \
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\"
RT2=\$(curl -s -X POST \"\$K1\$TOKEN_EP\" \
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\" \
| jget refresh_token)
# ── [6] 무효화가 반대 방향으로도 전파되는가 ───────────────────────────────
# keycloak-1 에서 로그아웃시키고, keycloak-0 에서 갱신을 시도한다.
echo '###STEP6_LOGOUT_VIA_K1'
curl -s -o /dev/null -w 'http_code=%{http_code}\n' -X POST \"\$K1/realms/master/protocol/openid-connect/logout\" \
-d client_id=admin-cli -d \"refresh_token=\$RT2\"
echo '###STEP7_REFRESH_ON_K0_AFTER_LOGOUT'
curl -s -w '\nhttp_code=%{http_code}\n' -X POST \"\$K0\$TOKEN_EP\" \
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT2\"
echo '###END'
" > "$OUT/raw.txt" 2>&1
sed -i '/^pod .* deleted$/d' "$OUT/raw.txt"
python3 - "$OUT/raw.txt" <<'PY' | tee "$OUT/report.txt"
import base64, json, sys
raw = open(sys.argv[1]).read()
blocks, cur = {}, None
for line in raw.splitlines():
if line.startswith('###'):
cur = line[3:]; blocks[cur] = []
elif cur is not None:
blocks[cur].append(line)
get = lambda k: '\n'.join(blocks.get(k, [])).strip()
def j(s):
try: return json.JSONDecoder().raw_decode(s.strip())[0]
except Exception: return None
def claims(tok):
p = tok.split('.')[1]; p += '=' * (-len(p) % 4)
return json.loads(base64.urlsafe_b64decode(p))
login = j(get('STEP1_LOGIN'))
if not login or 'access_token' not in login:
print('로그인 실패:', get('STEP1_LOGIN')[:300]); sys.exit(1)
ac = claims(login['access_token'])
rc = claims(login['refresh_token'])
SID = ac['sid']
print('=== [1] keycloak-0 에서 로그인 ===')
print(f" sid {SID}")
print(f" sub {ac.get('sub')}")
print(f" iss {ac.get('iss')}")
print(f" access 수명 {ac['exp']-ac['iat']}초")
print(f" refresh 수명 {rc['exp']-rc['iat']}초 typ={rc.get('typ')}")
print(f" refresh jti {rc.get('jti')}")
print()
print('=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===')
for step, who in (('STEP3_SESSIONS_K0', 'keycloak-0 (발급 노드)'),
('STEP3_SESSIONS_K1', 'keycloak-1 (반대편)')):
d = j(get(step))
if d is None:
print(f' {who:24} 파싱 실패: {get(step)[:120]}'); continue
ids = [s.get('id') for s in d]
mark = '보임 ✔' if SID in ids else '없음 ✘'
print(f' {who:24} 세션 {len(ids)}개 중 대상 sid → {mark}')
for s in d:
if s.get('id') == SID:
print(f" ipAddress={s.get('ipAddress')} start={s.get('start')} lastAccess={s.get('lastAccess')}")
def show(step, title, expect):
print(); print(f'=== {title} ===')
body = get(step)
code = [l for l in body.splitlines() if l.startswith('http_code=')]
code = code[0].split('=')[1] if code else '?'
d = j(body)
ok = '기대대로' if code == expect else f'기대({expect})와 다름'
print(f' HTTP {code} ← {ok}')
if d and 'access_token' in d:
c = claims(d['access_token'])
same = '동일 ✔' if c.get('sid') == SID else f"다름 ✘ ({c.get('sid')})"
print(f' 새 토큰의 sid → {same}')
elif d:
print(f" error {d.get('error')}")
print(f" error_description {d.get('error_description')}")
show('STEP5_REFRESH_ON_K1',
'[5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용', '200')
print(); print('=== [6] keycloak-1 을 통해 로그아웃 ===')
print(' ' + get('STEP6_LOGOUT_VIA_K1').strip())
show('STEP7_REFRESH_ON_K0_AFTER_LOGOUT',
'[7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파)', '400')
open('/tmp/session-replication/sid.txt','w').write(SID)
PY
SID=$(cat /tmp/session-replication/sid.txt 2>/dev/null)
echo
echo "=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ==="
echo " 대상 sid: $SID"
$PSQL "select user_session_id, offline_flag, created_on, last_session_refresh
from offline_user_session where user_session_id='$SID'" 2>/dev/null \
| sed 's/^/ /' | grep -q . \
&& $PSQL "select user_session_id||' | flag='||offline_flag||' | created='||created_on||' | refresh='||last_session_refresh
from offline_user_session where user_session_id='$SID'" 2>/dev/null | sed 's/^/ /' \
|| echo " 행 없음 — 로그아웃으로 삭제되었다"
echo
echo " 전체 세션 수: $($PSQL 'select count(*) from offline_user_session' 2>/dev/null)"
@@ -0,0 +1,13 @@
=== [기준선 1] 클러스터 뷰 — 양쪽 파드의 마지막 ISPN000094 ===
keycloak-0: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
keycloak-1: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
=== [기준선 2] JGROUPS_PING — 디스커버리 등록 ===
name | ip | coord | coordinated_by
------------------+-----------------+-------+---------------------------------------------
keycloak-0-30843 | 10.42.1.43:7800 | f | uuid://00000000-0000-0000-0000-000000000007
keycloak-1-48749 | 10.42.0.35:7800 | t | uuid://00000000-0000-0000-0000-000000000007
(2 rows)
=== [기준선 3] 기존 NetworkPolicy ===
No resources found in keycloak-lab namespace.
@@ -0,0 +1,17 @@
=== [대조군] 차단 전 — keycloak-0 로그인 → keycloak-1 에서 refresh ===
sid tAWs2gCPr6SOcD4jDR9-_CzB
keycloak-1 에서 refresh: 200
=== [기준선 4] JGroups 지표 — 양쪽 노드 ===
--- K0 (10.42.1.43) ---
vendor_jgroups_stats_bytes_sent_total 31476.0
vendor_jgroups_merge3_get_num_merge_events 0.0
vendor_jgroups_merge3_get_views 0.0
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
--- K1 (10.42.0.35) ---
vendor_jgroups_merge3_get_views 0.0
vendor_jgroups_stats_bytes_sent_total 126765.0
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
vendor_jgroups_merge3_get_num_merge_events 0.0
@@ -0,0 +1,8 @@
=== 차단 적용 ===
networkpolicy.networking.k8s.io/a1-block-jgroups-transport created
a1-block-jgroups-transport map[app:keycloak]
적용 시각: 11:38:08
=== FD_SOCK2 가 상대를 의심하기까지 기다린다 (15초 간격, 최대 3분) ===
+15초 suspected(k0 k1) =
→ 변화 감지
@@ -0,0 +1,26 @@
차단 경과: 11:38:51 (적용 11:38:08)
=== [차단 후 1] JGroups 지표 ===
--- keycloak-0 ---
vendor_jgroups_stats_bytes_sent_total 32857.0
vendor_jgroups_merge3_get_num_merge_events 0.0
vendor_jgroups_merge3_get_views 0.0
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
--- keycloak-1 ---
vendor_jgroups_merge3_get_views 0.0
vendor_jgroups_stats_bytes_sent_total 129103.0
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
vendor_jgroups_merge3_get_num_merge_events 0.0
=== [차단 후 2] 클러스터 뷰 — 갈라졌는가 ===
keycloak-0:
keycloak-1:
=== [차단 후 3] JGROUPS_PING — 디스커버리는 살아 있는가 ===
name | ip | coord
------------------+-----------------+-------
keycloak-0-30843 | 10.42.1.43:7800 | f
keycloak-1-48749 | 10.42.0.35:7800 | t
(2 rows)
@@ -0,0 +1,16 @@
=== [문제 확정] NetworkPolicy 적용 후에도 기존 연결이 conntrack 에 살아 있다 ===
--- kc-lab-1 ---
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 src=10.42.1.43 dst=10.42.0.35 sport=57800 dport=50477 [ASSURED] mark=0 use=1
--- kc-lab-2 ---
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1
tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED] src=10.42.0.35 dst=10.42.1.58 sport=7800 dport=34824 mark=0 use=1
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 src=10.42.1.43 dst=10.42.0.35 sport=57800 dport=50477 [ASSURED] mark=0 use=1
=== [조치] 7800 흐름의 conntrack 항목을 지운다 → 다음 패킷이 정책을 다시 탄다 ===
kc-lab-1: tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1 conntrack v1.4.7 (conntrack-tools): 0 flow entries have been deleted.
kc-lab-2: tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED] src=10.42.0.35 dst=10.42.1.58 sport=7800 dport=34824 mark=0 use=1 conntrack v1.4.7 (conntrack-tools): 0 flow entries have been deleted.
=== 삭제 후 7800 conntrack ===
kc-lab-1: 2 건
kc-lab-2: 2 건
@@ -0,0 +1,13 @@
관찰 시작: 11:42:03
+20초 suspected(k0 k1) = []
+40초 suspected(k0 k1) = []
+60초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
+80초 suspected(k0 k1) = []
+100초 suspected(k0 k1) = []
+120초 suspected(k0 k1) = []
+140초 suspected(k0 k1) = [0.0 ]
+160초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
=== 클러스터 뷰 변화 (최근 8분) ===
--- keycloak-0 ---
--- keycloak-1 ---
@@ -0,0 +1,16 @@
=== vendor_cluster_size — 지난 25분 (차단 11:38:08, conntrack 삭제 11:41) ===
keycloak-0:
11:20=2 11:21=2 11:22=2 11:23=2 11:24=2 11:25=2 11:26=2 11:27=2 11:28=2 11:29=2 11:30=2 11:31=2 11:32=2 11:33=2 11:34=2 11:35=2 11:36=2 11:37=2 11:38=2 11:39=2 11:40=2 11:41=2 11:42=2 11:43=2 11:44=2 11:45=2
keycloak-1:
11:20=2 11:21=2 11:22=2 11:23=2 11:24=2 11:25=2 11:26=2 11:27=2 11:28=2 11:29=2 11:30=2 11:31=2 11:32=2 11:33=2 11:34=2 11:35=2 11:36=2 11:37=2 11:38=2 11:39=2 11:40=2 11:41=2 11:42=2 11:43=2 11:44=2 11:45=2
=== 현재 값 ===
keycloak-1 = 2 멤버
keycloak-0 = 2 멤버
=== 7800 소켓 상태 (파드 내부) ===
keycloak-0 2
keycloak-1 2
=== conntrack ===
kc-lab-1 1 건
kc-lab-2 1 건
@@ -0,0 +1,9 @@
=== 정책이 걸린 상태에서 keycloak-0 을 재시작한다 → 재연결이 막힌다 ===
재시작 시각: 11:46:07
pod "keycloak-0" deleted from keycloak-lab namespace
keycloak-0 false 10.42.1.67 2026-09-04T02:44:23Z
=== cluster_size 추이 ===
keycloak-0: 11:45:27=1 11:45:57=1 11:46:27=1 11:46:57=1 11:47:27=1
keycloak-0: 11:40:57=2 11:41:27=2 11:41:57=2 11:42:27=2 11:42:57=2 11:43:27=2 11:43:57=2
keycloak-1: 11:40:57=2 11:41:27=2 11:41:57=2 11:42:27=2 11:42:57=2 11:43:27=2 11:43:57=2 11:44:27=1 11:44:57=1 11:45:27=1 11:45:57=1 11:46:27=1 11:46:57=1 11:47:27=1
@@ -0,0 +1,16 @@
=== keycloak-0 헬스 상태 ===
keycloak-0 = 10.42.1.67 keycloak-1 = 10.42.0.35
PodReadyToStartContainers=True
Initialized=True
Ready=False ContainersNotReady
ContainersReady=False ContainersNotReady
PodScheduled=True
=== ★ 본 시험 — 분단 상태에서 교차 노드 세션이 되는가 ===
[1] keycloak-0 로그인 sid=nShl5TaBrZnKStDqaspjgmJB
[2] keycloak-1 에서 refresh HTTP 200
[3] keycloak-1 에서 로그아웃 HTTP 204
[4] keycloak-0 에서 재갱신 시도 HTTP 200
(400 이면 무효화가 전파된 것)
=== DB 세션 수 ===
@@ -0,0 +1,33 @@
=== 그 sid 가 DB 에 남아 있는가 ===
user_session_id | offline_flag | last_session_refresh
-----------------+--------------+----------------------
(0 rows)
=== 전체 온라인 세션 수 ===
1
=== 노드별 세션 캐시 엔트리 (Prometheus) ===
keycloak-1 kc-lab-1 = 0
keycloak-0 kc-lab-2 = 1
=== keycloak-0 이 Ready 가 아닌 이유 — 헬스 응답 ===
{
"status": "DOWN",
"checks": [
{
"name": "Graceful Shutdown",
"status": "UP"
},
{
"name": "Keycloak cluster health check",
"status": "DOWN",
"data": {
"Failing since": "2026-09-04 02:45:14,251"
}
},
{
"name": "Keycloak database connections async health check",
"status": "UP"
},
{
"name": "Keycloak Initialized",
@@ -0,0 +1,25 @@
=== 양쪽 노드의 readiness — 둘 다 DOWN 이면 전면 장애다 ===
Traceback (most recent call last):
File "<string>", line 3, in <module>
d=json.load(sys.stdin)
File "/usr/lib/python3.14/json/__init__.py", line 298, in load
return loads(fp.read(),
cls=cls, object_hook=object_hook,
parse_float=parse_float, parse_int=parse_int,
parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw)
File "/usr/lib/python3.14/json/__init__.py", line 352, in loads
return _default_decoder.decode(s)
~~~~~~~~~~~~~~~~~~~~~~~^^^
File "/usr/lib/python3.14/json/decoder.py", line 348, in decode
raise JSONDecodeError("Extra data", s, end)
json.decoder.JSONDecodeError: Extra data: line 21 column 2 (char 446)
=== 파드 Ready 상태 ===
keycloak-0 false 0
keycloak-1 true 0
=== ★ Service 엔드포인트 — 트래픽을 받는 파드가 남아 있는가 ===
ready 주소: [10.42.0.35] notReady : [10.42.1.67]
=== ★ 외부 진입점으로 실제 로그인이 되는가 (nginx→Traefik→Service) ===
https://auth.hyeonworks.com/realms/master HTTP 200
토큰 발급 HTTP 200
@@ -0,0 +1,24 @@
=== 차단 해제 ===
해제 시각: 11:49:58
networkpolicy.networking.k8s.io "a1-block-jgroups-transport" deleted from keycloak-lab namespace
=== 자동으로 다시 붙는가 (30초 간격, 최대 4분) ===
+30초 keycloak-0=1 keycloak-1=1 | Ready 파드 2 개
+60초 keycloak-0=1 keycloak-1=1 | Ready 파드 2 개
+90초 keycloak-0=2 keycloak-1=2 | Ready 파드 3 개
→ 클러스터 재형성
=== 복구 로그 ===
keycloak-0: [keycloak-0-26403(v=16.0.12)|0] (1) [keycloak-0-26403(v=16.0.12)]
keycloak-1: [keycloak-1-48749(v=16.0.12)|6] (1) [keycloak-1-48749(v=16.0.12)]
=== MERGE3 가 합쳤는가 ===
merge_events keycloak-1 = 1
merge_events keycloak-0 = 1
=== JGROUPS_PING — 코디네이터가 하나로 돌아왔는가 ===
name | ip | coord
------------------+-----------------+-------
keycloak-0-26403 | 10.42.1.67:7800 | t
keycloak-1-48749 | 10.42.0.35:7800 | f
(2 rows)
@@ -0,0 +1,26 @@
# A-1 — JGroups 트랜스포트(7800) 차단 증거
2026-09-04 11:3811:52 KST · Keycloak 26.7.0 / Infinispan 16.0.12
해설: [`docs/experiment-a1-jgroups-transport-block.md`](../../experiment-a1-jgroups-transport-block.md)
| 파일 | 무엇을 보여주는가 |
|---|---|
| `01-baseline-cluster.txt` | 차단 전 — 양쪽이 뷰 ID 5·멤버 2로 일치, `JGROUPS_PING` 코디네이터 1명 |
| `02-control-before-block.txt` | **대조군** — 차단 전 교차 노드 refresh `200`, JGroups 지표 전부 0 |
| `03-block-applied.txt` | NetworkPolicy 적용. **빈 측정값을 "변화 감지"로 오판한 기록** |
| `04-after-block-state.txt` | 차단 43초 후 — 지표 무변화, `JGROUPS_PING` 그대로 |
| `05-conntrack-problem.txt` | **핵심 문제**`ESTABLISHED [ASSURED]` 로 기존 연결이 살아 있음. FD_SOCK2 의 **57800** 포트도 함께 드러남 |
| `06-partition-observed.txt` | 임시 curl 파드 폴링의 실패 — 빈 값·개수 불일치 |
| `07-cluster-size.txt` | **`vendor_cluster_size` 가 25분 내내 2** — 분단이 일어나지 않았다는 결정적 증거 |
| `08-restart-forced-partition.txt` | 재연결 강제 후 `2 → 1` |
| `09-cross-node-under-partition.txt` | **본 시험** — 교차 refresh `200`(예측 적중), **로그아웃 후 재갱신 `200`(예측 빗나감)** |
| `10-logout-not-propagated.txt` | 기제 확정 — **DB 행 0건인데 keycloak-0 캐시에 1건**, 헬스체크 `cluster health: DOWN` |
| `11-service-impact.txt` | **분단 노드가 Service 에서 빠짐.** `ready=[10.42.0.35] notReady=[10.42.1.67]`, 외부 로그인 `200` |
| `12-recovery.txt` | 90초 만에 자동 재형성, `merge3_get_num_merge_events = 1`, 코디네이터 재선출 |
| `a1-cluster-size-partition-recovery.png` | Grafana — `vendor_cluster_size``2 → 1 → 2` 로 움직이는 전 구간 |
## 핵심 세 줄
1. **NetworkPolicy 만으로는 이미 붙어 있는 클러스터를 못 끊는다.** conntrack 의 ESTABLISHED 가 먼저 통과시킨다.
2. **세션 공유는 분단을 견딘다(200).** 통념이 틀렸고 A-0 모델이 맞다.
3. **로그아웃 무효화는 7800 을 탄다.** DB 행이 지워져도 반대편은 낡은 캐시로 200 을 준다 — A-0 의 인과 해석을 정정한다.
Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

@@ -0,0 +1,14 @@
=== A-2 기준선 — 클러스터가 정상으로 돌아왔는가 ===
keycloak-0 true 10.42.1.67 kc-lab-2
keycloak-1 true 10.42.0.35 kc-lab-1
postgres-7b474b88c8-sn9ff true 10.42.1.24 kc-lab-2
cluster_size keycloak-1 = 2
cluster_size keycloak-0 = 2
=== 노드별 세션 캐시 (실험 설계에 필요) ===
keycloak-1 kc-lab-1 = 0 건
keycloak-0 kc-lab-2 = 0 건
=== DB 온라인 세션 ===
2
@@ -0,0 +1,10 @@
pod/a2-probe condition met
keycloak-0=10.42.1.67 keycloak-1=10.42.0.35
=== [준비] 양쪽 노드에 세션을 하나씩 만든다 ===
keycloak-0 에서 로그인 sid=EAXV5HcG2J1BZ3vnwONf64AQ 토큰길이=613
keycloak-1 에서 로그인 sid=McyTj5lj3n_JqApCXeuAHExc 토큰길이=613
=== [확인] 세션이 각자 노드에만 캐시되었는가 ===
keycloak-1 = 0 건
keycloak-0 = 1 건
@@ -0,0 +1,16 @@
=== [1] 토큰을 새로 발급 (access 수명 60초) ===
발급 완료 sid=RKXQGAgkuLtouFMVPTFmp_0_
=== [2] PostgreSQL 정지 ===
정지 시각: 11:56:04
deployment.apps/postgres scaled
pod/postgres-7b474b88c8-sn9ff condition met
삭제 완료: 11:56:04
=== [3] 네 경로를 즉시 시험 ===
④ 이미 발급된 access token 으로 관리 API HTTP 000000{"error":"HTTP 401 Unauthorized"}401
① 캐시를 가진 노드(keycloak-0)에서 refresh HTTP 500
② 캐시가 없는 노드(keycloak-1)에서 refresh HTTP 500
③ 새 로그인 HTTP 500
--- 오류 본문 (새 로그인) ---
{"error":"unknown_error","error_description":"For more on this error consult the server log."}
@@ -0,0 +1,29 @@
=== 파드 Ready 상태 — DB 가 없으면 어떻게 되는가 ===
keycloak-0 false 0
keycloak-1 false 0
=== Service 엔드포인트 ===
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
notReady: [10.42.0.35 10.42.1.67]
=== health/ready 상세 ===
전체: DOWN
Graceful Shutdown UP
Keycloak cluster health check UP
Keycloak database connections async health check DOWN
Keycloak Initialized UP
=== ④ 다시 — 서명 검증만 필요한 경로는 살아 있는가 ===
JWKS 엔드포인트(realm 공개키) HTTP 200
realm 메타데이터(.well-known) HTTP 200
관리 API(세션 조회 필요) HTTP 500
=== 외부 진입점 ===
https://auth.hyeonworks.com/realms/master HTTP 503
=== Keycloak 로그 — 실제 오류 ===
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call(ConnectionPool.java:664)
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call(ConnectionPool.java:645)
Caused by: java.net.ConnectException: Connection refused
at org.postgresql.core.v3.ConnectionFactoryImpl.tryConnect(ConnectionFactoryImpl.java:219)
at org.postgresql.core.v3.ConnectionFactoryImpl.openConnectionImpl(ConnectionFactoryImpl.java:365)
@@ -0,0 +1,21 @@
=== ★ up 지표는 무엇을 말하는가 (프로세스는 살아 있다) ===
up{pod=keycloak-1} = 1 ← 1 인데 서비스는 503 이다
up{pod=keycloak-0} = 1 ← 1 인데 서비스는 503 이다
=== 복구 — PostgreSQL 재기동 ===
재기동 시각: 11:57:09
deployment.apps/postgres scaled
Waiting for deployment "postgres" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "postgres" rollout to finish: 0 of 1 updated replicas are available...
deployment "postgres" successfully rolled out
=== Keycloak 이 스스로 회복하는가 (재시작 없이) ===
+15초 keycloak-0 true keycloak-1 true | 외부 HTTP 200
→ 서비스 복귀
=== 재시작 횟수 — 파드가 죽었다 살아난 것인가, 그대로 회복한 것인가 ===
keycloak-0 0
keycloak-1 0
=== 정지 전 세션이 살아남았는가 ===
online 세션 5
+19
View File
@@ -0,0 +1,19 @@
# A-2 — PostgreSQL 정지 증거
2026-09-04 11:5611:58 KST · Keycloak 26.7.0
해설: [`docs/experiment-a2-database-loss.md`](../../experiment-a2-database-loss.md)
| 파일 | 무엇을 보여주는가 |
|---|---|
| `01-baseline.txt` | 정지 전 — 양쪽 Ready, `cluster_size=2` |
| `02-setup-sessions.txt` | 양쪽 노드에 세션 하나씩. 캐시는 각자 노드에만 |
| `03-four-paths.txt` | **네 경로 전부 `500`.** 캐시를 가진 노드도 실패 — refresh 는 쓰기다 |
| `04-health-and-service.txt` | **전면 장애 증거** — Ready 파드 0개, `ready 주소=[]`, 외부 **503**, `database connections: DOWN`. JWKS·.well-known 은 `200` |
| `05-recovery.txt` | **`up=1` 인 채로 503.** DB 복귀 15초 후 재시작 0회로 자동 회복, 세션 5건 생존 |
| `a2-up-stayed-1-during-outage.png` | Grafana — `up{job="keycloak"}` 이 전면 장애 내내 **1에 평평** |
## 핵심 세 줄
1. **DB 는 단일 장애점이다.** Keycloak 을 몇 대로 늘려도 같이 죽는다 — Ready 파드 0개, 외부 503.
2. **캐시는 읽기를 대신할 뿐 쓰기를 못 한다.** refresh 는 `UPDATE LAST_SESSION_REFRESH` 를 하므로 캐시가 있어도 실패한다.
3. **`up` 은 이 장애를 못 잡는다.** 알림은 readiness 와 외부 응답 코드에 걸어야 한다.
Binary file not shown.

After

Width:  |  Height:  |  Size: 60 KiB

@@ -0,0 +1,19 @@
=== [준비] 손실 측정 설계 확인 ===
LAST_SESSION_REFRESH 는 integer(초) — 200ms 손실은 보이지 않는다
created_on | integer | | not null |
last_session_refresh | integer | | not null | 0
"idx_user_session_expiration_created" btree (realm_id, offline_flag, remember_me, created_on, user_session_id, user_id)
"idx_user_session_expiration_last_refresh" btree (realm_id, offline_flag, remember_me, last_session_refresh, user_session_id, user_id)
→ 대신 행 존재 여부로 잰다. 로그인 하나 = 행 하나 = 이진 판정
전역 synchronous_commit: on
=== [1] 빠른 연속 로그인을 백그라운드로 시작 ===
루프 시작
6초 경과 — 지금까지 성공한 로그인: 0
=== [2] PostgreSQL 강제 종료 (SIGKILL) ===
종료 시각: 12:00:26.511
pod "postgres-7b474b88c8-xc2vt" force deleted from keycloak-lab namespace
삭제 반환: 12:00:26.586
클라이언트가 200 을 받은 로그인 수: 0
@@ -0,0 +1,14 @@
deployment "postgres" successfully rolled out
=== crash recovery 가 실행되었는가 (강제 종료의 흔적) ===
2026-09-04 02:58:41.036 UTC [1] LOG: database system is ready to accept connections
=== [설계 확인] 로그인 트랜잭션도 synchronous_commit 을 끄는가 ===
--- 로그인 트랜잭션 (INSERT 가 있는 것) ---
2:BEGIN
5:COMMIT
6:BEGIN
9:insert into OFFLINE_USER_SESSION (BROKER_SESSION_ID,CREATED_ON,DATA,LAST_SESSION_REFRESH,REALM_ID,REMEMBER_ME,USER_ID,VERSION,OFFLINE_FLAG,USER_SESSION_ID) values ($1,$2,$3,$4,$5,$6,$7,$8,$9,$10)
10:insert into OFFLINE_CLIENT_SESSION (DATA,REALM_ID,TIMESTAMP,VERSION,CLIENT_ID,CLIENT_STORAGE_PROVIDER,EXTERNAL_CLIENT_ID,OFFLINE_FLAG,USER_SESSION_ID) values ($1,$2,$3,$4,$5,$6,$7,$8,$9)
11:SET LOCAL synchronous_commit TO OFF
12:COMMIT
@@ -0,0 +1,17 @@
=== [1] 로그인 루프 시작 (호스트에서 백그라운드로 exec — 세션이 살아 있어야 한다) ===
8초 동안 클라이언트가 200 을 받은 로그인: 106 건
=== [2] SIGKILL ===
종료: 12:01:32.981
반환: 12:01:33.236
최종 성공 로그인 수: 110 건
마지막 sid: FimM-krSybBACP2qIvshLWwU
마지막 sid: EwFfFwOIfqiv8N5GQ5OjtsVq
마지막 sid: CJX-PxFkS7rUc_9FQgB7iw1f
마지막 sid: 1EFK7SgUA4M7tq_SkC_BD2er
마지막 sid: _LiqTczuyxlpOs3T3xs25SLv
=== [3] PostgreSQL 재기동 후 crash recovery 확인 ===
Waiting for deployment "postgres" rollout to finish: 0 of 1 updated replicas are available...
deployment "postgres" successfully rolled out
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
@@ -0,0 +1,27 @@
=== [4] 클라이언트가 받은 sid 가 DB 에 있는가 ===
클라이언트가 200 을 받은 sid: 291 건
DB 온라인 세션 총계: 375
--- 마지막 15건을 하나씩 조회 ---
TCCOYnVlN30Y2fGyJEsVJ_Lq 있음
vBcllIkKWovh-FXN9tSzmxAe 있음
eMpN_ywUdRTks9uBE9aTumPK 있음
aPC_T0yrlMjrlskcLAp0AvY6 있음
UBPxmduB-ahGHBg636sY3AEz 있음
HLnBloNX9R3qQJSkpOnkNqL4 있음
_KPJS30IAqVhkTJGHcCxKvrM 있음
rq3caZ9MkyMYlFSSzRjQLygD 있음
ivvQm70hjl55DPpF7vpYmz_E 있음
81mx-rmi-tAeogHx3-su3z2q 있음
WnvNDH93uzcz1XNFbaMSXk1A 있음
DXPAIhjO5sGpCUlcD8IEoS8L 있음
aZMvl4IwPdK-rC_7bG005Z5C 있음
eIuBCprfWA5x0glcgSKYrrX0 있음
ozES5kEeu2IFf_cfcC_jFlbF 있음
마지막 15건 중 유실: 0 건
=== [5] 전체 대조 — 몇 건이나 사라졌는가 ===
클라이언트 성공: 291 건
DB 에 존재: 291 건
★ 유실: 0 건
@@ -0,0 +1,12 @@
=== [정리] 세션 테이블 비우고 루프 잔여 확인 ===
DELETE 375
남은 세션: 0
=== [재주입] postmaster(PID 1)에 SIGKILL — 진짜 크래시 ===
8초 후 성공 로그인: 110 건
SIGKILL: 12:03:21.441
최종 성공 로그인: 139 건
=== [검증] 이번엔 crash recovery 가 돌았는가 ===
deployment "postgres" successfully rolled out
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
@@ -0,0 +1,17 @@
=== 로그인 루프 시작 ===
8초 후: 112 건
=== 백엔드 프로세스에 SIGKILL → postmaster 가 재초기화한다 ===
시각: 12:04:22.063
최종 성공 로그인: 153 건
=== [검증] crash recovery 가 돌았는가 ===
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
2026-09-04 03:02:35.807 UTC [1] LOG: server process (PID 40) was terminated by signal 9: Killed
2026-09-04 03:02:35.807 UTC [1] LOG: terminating any other active server processes
2026-09-04 03:02:35.814 UTC [1] LOG: all server processes terminated; reinitializing
2026-09-04 03:02:35.896 UTC [2585] LOG: database system was not properly shut down; automatic recovery in progress
2026-09-04 03:02:35.899 UTC [2585] LOG: redo starts at 0/23CAB68
2026-09-04 03:02:35.904 UTC [2585] LOG: redo done at 0/2529E40 system usage: CPU: user: 0.00 s, system: 0.00 s, elapsed: 0.00 s
2026-09-04 03:02:35.923 UTC [2586] LOG: checkpoint complete: wrote 113 buffers (0.7%); 0 WAL file(s) added, 0 removed, 0 recycled; write=0.004 s, sync=0.004 s, total=0.015 s; sync files=27, longest=0.003 s, average=0.001 s; distance=1405 kB, estimate=1405 kB; lsn=0/252A048, redo lsn=0/252A048
2026-09-04 03:02:35.926 UTC [1] LOG: database system is ready to accept connections
@@ -0,0 +1,19 @@
=== 크래시 전후 대조 ===
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
그중 DB 에 실제로 존재 : 149 건
★ 유실 : 4 건
DB 전체 온라인 세션 : 150 건
=== 유실된 sid 목록 ===
★ CQUfg9HLH29xvhiu6pVlfWOo ← 토큰은 발급됐는데 세션이 없다
★ 5gLP4fqmpZBbjhH_d-0TPMMr ← 토큰은 발급됐는데 세션이 없다
★ hkcOv1QskUFmYveMLB6Hljra ← 토큰은 발급됐는데 세션이 없다
★ p5XybeQIYmAs818gO4Vl_5ea ← 토큰은 발급됐는데 세션이 없다
=== 그 토큰이 지금 실제로 쓰이는가 (마지막 sid 로 확인) ===
마지막 sid: 8do0Bw6tkVLDVxgxotE7GosH
user_session_id | created_on | last_session_refresh
--------------------------+------------+----------------------
8do0Bw6tkVLDVxgxotE7GosH | 1788490958 | 1788490958
(1 row)
+20
View File
@@ -0,0 +1,20 @@
# A-3 — DB 강제 종료와 데이터 손실 증거
2026-09-04 12:0012:05 KST · Keycloak 26.7.0 / PostgreSQL 16
해설: [`docs/experiment-a3-database-crash.md`](../../experiment-a3-database-crash.md)
| 파일 | 무엇을 보여주는가 |
|---|---|
| `01-crash-injection.txt` | 첫 시도 실패 — 파드 안 백그라운드 루프가 `exec` 종료와 함께 죽어 0건 수집 |
| `02-design-check.txt` | **핵심 설계 확인** — 로그인 트랜잭션도 `SET LOCAL synchronous_commit TO OFF` 로 커밋한다 |
| `03-loss-measurement.txt` | `--grace-period=0 --force` 주입 |
| `04-comparison.txt` | **유실 0건** — 그러나 crash recovery 가 안 돌았다. 죽인 적이 없는 것 |
| `05-true-crash.txt` | `kill -9 1` 시도 — **컨테이너 안에서 PID 1 은 SIGKILL 을 무시한다** |
| `06-backend-kill-crash.txt` | **성공한 주입** — 백엔드에 SIGKILL → `not properly shut down` / `redo starts` / `redo done` |
| `07-loss-result.txt` | **결과: 153건 중 4건 유실.** 토큰은 발급됐는데 세션 행이 없는 sid 목록 |
## 핵심 세 줄
1. **로그인도 비동기 커밋이다.** refresh 시각뿐 아니라 **로그인 자체**가 사라질 수 있다.
2. **153건 중 4건(약 2.6%) 유실** — 초당 19건 기준 마지막 0.2초 분량, `wal_writer_delay` 기본값과 일치.
3. **주입을 세 번 시도해 세 번째에 성공했다.** 앞의 둘은 "손실 0"으로 보였지만 실제로는 크래시가 아니었다.
@@ -0,0 +1,56 @@
수집 시각: 2026-09-03 17:24:54 KST
대상: Keycloak 26.7.0 × 2 + PostgreSQL 16, k3s 2노드
=== [1] 파드 배치 ===
keycloak-0 1/1 10.42.1.18 kc-lab-2
keycloak-1 1/1 10.42.0.16 kc-lab-1
postgres-7b474b88c8-bw7b8 1/1 10.42.1.19 kc-lab-2
=== [2] 클러스터 뷰 로그 (Infinispan) ===
-- keycloak-0 --
2026-09-03 08:18:23,359 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938(v=16.0.12), keycloak-0-49501(v=16.0.12)]
2026-09-03 08:18:23,433 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000079: Channel `ISPN` local address is `keycloak-0-49501`, physical addresses are `[10.42.1.18:7800]`
-- keycloak-1 --
2026-09-03 08:18:23,269 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938(v=16.0.12), keycloak-0-49501(v=16.0.12)]
2026-09-03 08:18:23,282 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN100000: Node keycloak-0-49501 joined the cluster
2026-09-03 08:18:23,286 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN100000: Node keycloak-0-49501 joined the cluster
=== [3] JGROUPS_PING 테이블 구조 ===
Table "public.jgroups_ping"
Column | Type | Collation | Nullable | Default
----------------+------------------------+-----------+----------+---------
address | character varying(200) | | not null |
name | character varying(200) | | |
cluster_name | character varying(200) | | not null |
ip | character varying(200) | | not null |
coord | boolean | | |
last_update | bigint | | |
coordinated_by | character varying(200) | | |
Indexes:
"constraint_jgroups_ping" PRIMARY KEY, btree (address)
=== [4] JGROUPS_PING 등록 내역 ===
name | cluster_name | ip | coord
------------------+--------------+-----------------+-------
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
(2 rows)
=== [5] 외부 접근 — OIDC discovery ===
issuer https://auth.hyeonworks.com/realms/master
authorization_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/auth
token_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/token
end_session_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/logout
jwks_uri https://auth.hyeonworks.com/realms/master/protocol/openid-connect/certs
★ 전부 https. 첫 실험에서 확정한 KC_HOSTNAME + KC_PROXY_HEADERS 조합이 작동한다.
=== [6] 자원 사용 ===
keycloak-0 8m 594Mi
keycloak-1 9m 593Mi
postgres-7b474b88c8-bw7b8 3m 67Mi
--- 노드 ---
kc-lab-1 2248Mi (65%)
kc-lab-2 1447Mi (58%)
@@ -0,0 +1,62 @@
# 증거 — Keycloak 멀티노드 클러스터 형성
`docs/keycloak-multinode-cluster.md`의 근거 자료.
**정상적으로 클러스터가 형성된 상태**에서 수집했으며, 이후 고장을 주입한
뒤 이것과 대조한다.
수집 시각: 2026-09-03 17:24 KST
| 파일 | 내용 |
|---|---|
| `01-cluster-formed.txt` | 파드 배치·클러스터 뷰 로그·JGROUPS_PING·OIDC discovery·자원 |
## 이 상태에서 확인된 것
**클러스터 뷰가 멤버 2를 보고한다**
```
ISPN000094: Received new cluster view for channel ISPN:
[keycloak-1-26938|1] (2) [keycloak-1-26938, keycloak-0-49501]
ISPN100000: Node keycloak-0-49501 joined the cluster
ISPN000079: physical addresses are [10.42.1.18:7800]
```
**디스커버리와 통신 경로가 한 테이블에 다 보인다**
```
name | cluster_name | ip | coord
------------------+--------------+-----------------+-------
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
```
`name`/`cluster_name`은 **DB 디스커버리**의 결과이고, `ip``:7800`
**실제 통신 경로**다. 7800을 막으면 이 표는 그대로 채워지면서 클러스터 뷰만
깨질 것으로 예상한다 — 다음 실험의 가설이다.
`coord = t``keycloak-1`이 코디네이터다.
**배치** — 서로 다른 노드에 하나씩. PostgreSQL은 `kc-lab-2`에 있으므로
**그 노드를 죽이면 Keycloak 하나와 DB가 동시에 사라진다.**
```
keycloak-0 10.42.1.18 kc-lab-2
keycloak-1 10.42.0.16 kc-lab-1
postgres 10.42.1.19 kc-lab-2
```
**issuer가 https로 발급된다** — 첫 실험(2홉 헤더 계약)의 결론이 적용된 결과다.
## 재수집
```bash
kubectl -n keycloak-lab get pods -o wide
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
PG=$(kubectl -n keycloak-lab get pod -l app=postgres -o name | head -1)
kubectl -n keycloak-lab exec "$PG" -- \
psql -U keycloak -d keycloak -c "SELECT name, cluster_name, ip, coord FROM jgroups_ping ORDER BY name;"
curl -s https://auth.hyeonworks.com/realms/master/.well-known/openid-configuration | python3 -m json.tool
kubectl -n keycloak-lab top pods
```
@@ -0,0 +1,52 @@
===================================================================
실험 0 — 한 노드에서 만든 세션이 다른 노드에서 쓰이는가
===================================================================
### 사전 확인: 클러스터가 2 멤버로 형성되었는가
2026-09-04 00:52:09,294 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
name | ip | coord
------------------+-----------------+-------
keycloak-1-48749 | 10.42.0.35:7800 | t
keycloak-0-30843 | 10.42.1.43:7800 | f
(2 rows)
수집 시각: 2026-09-04 09:54:29 KST
=== 대상 ===
keycloak-0 10.42.1.43 kc-lab-2
keycloak-1 10.42.0.35 kc-lab-1
=== [0] 실험 전 DB 세션 ===
=== [1] keycloak-0 에서 로그인 ===
sid jiv3rVZi1VeaO07oVJkL_MYW
sub None
iss https://auth.hyeonworks.com/realms/master
access 수명 60초
refresh 수명 1800초 typ=Refresh
refresh jti 7669cc49-4778-851f-3c49-65f76964ae8e
=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
ipAddress=10.42.1.44 start=1788483164000 lastAccess=1788483164000
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
ipAddress=10.42.1.44 start=1788483164000 lastAccess=1788483164000
=== [5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용 ===
HTTP 200 ← 기대대로
새 토큰의 sid → 동일 ✔
=== [6] keycloak-1 을 통해 로그아웃 ===
http_code=204
=== [7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파) ===
HTTP 400 ← 기대대로
error invalid_grant
error_description Session not active
=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ===
대상 sid: jiv3rVZi1VeaO07oVJkL_MYW
행 없음 — 로그아웃으로 삭제되었다
전체 세션 수: 1
@@ -0,0 +1,35 @@
===================================================================
실험 0b — Infinispan 이 복제한 것인가, DB 를 같이 본 것인가
===================================================================
수집 시각: 2026-09-04 09:54:41 KST
=== 로그인은 keycloak-0 에만 보냈다 ===
로그인 응답: http_code=200
=== keycloak-0 (로그인을 받은 노드) ===
계수기 캐시 전 후 증가
rpc.replication_count clientSessions 1 1 +0
rpc.replication_count sessions 1 1 +0
approximate_entries_unique clientSessions 1 2 +1 ←
approximate_entries_unique sessions 1 2 +1 ←
hits clientSessions 2 2 +0
hits sessions 2 2 +0
misses clientSessions 2 3 +1 ←
misses sessions 3 4 +1 ←
stores clientSessions 2 3 +1 ←
stores sessions 2 3 +1 ←
=== keycloak-1 (아무 요청도 받지 않은 노드) ===
계수기 캐시 전 후 증가
rpc.replication_count clientSessions 7 7 +0
rpc.replication_count sessions 7 7 +0
approximate_entries_unique clientSessions 0 0 +0
approximate_entries_unique sessions 0 0 +0
hits clientSessions 4 4 +0
hits sessions 4 4 +0
misses clientSessions 0 0 +0
misses sessions 0 0 +0
stores clientSessions 1 1 +0
stores sessions 1 1 +0
@@ -0,0 +1,15 @@
===================================================================
실험 0c — 세션 엔트리는 어느 노드에 있는가 (로컬 캐시인가 분산인가)
===================================================================
수집 시각: 2026-09-04 09:54:54 KST
keycloak-0 = 10.42.1.43 (kc-lab-2)
keycloak-1 = 10.42.0.35 (kc-lab-1)
단계 k0 entries k1 entries
시작 2.0 0.0
keycloak-1 에 로그인 5회 2.0 5.0
keycloak-0 에 로그인 5회 7.0 5.0
=== 대조: PostgreSQL 에는 몇 건인가 ===
online 세션 12
@@ -0,0 +1,56 @@
===================================================================
실험 0d — 반대편 노드가 정말 DB 에서 읽는가 (SQL 을 직접 잡는다)
===================================================================
수집 시각: 2026-09-04 10:14:17 KST
keycloak-0 = 10.42.1.43 (세션을 만드는 노드)
keycloak-1 = 10.42.0.35 (읽기만 하는 노드)
=== PostgreSQL 문장 로깅을 켠다 ===
log_statement = all
log_line_prefix = %m [%p] %h
=== 요청 ===
SID=jSt9GEPVQLJsO-1CeJjVgltg
K1_ENTRIES_BEFORE=5.0
REFRESH_ON_K1=200
K1_ENTRIES_AFTER=5.0
=== PostgreSQL 문장 로깅을 끈다 ===
log_statement = none
=== keycloak-1 이 실제로 보낸 SQL 문장 ===
(파라미터가 $1 로 묶여 있어, sid 는 바로 아래 DETAIL 줄에 있다)
select puse1_0.OFFLINE_FLAG,puse1_0.USER_SESSION_ID,puse1_0.BROKER_SESSION_ID,puse1_0.CREATED_ON,puse1_0.DATA,puse1_0.LAST_SESSION_REFRESH,puse1_0.REALM_ID,puse1_0.REMEMBER_ME,puse1_0.USER_ID,puse1_0.VERSION from OFFLINE_USER_SESSION puse1_0 where (puse1_0.OFFLINE_FLAG,puse1_0.USER_SESSION_ID) in (($1,$2))
select puse1_0.VERSION from OFFLINE_USER_SESSION puse1_0 where puse1_0.USER_SESSION_ID=$1 and puse1_0.OFFLINE_FLAG=$2 for no key update of puse1_0 skip locked
select pcse1_0.CLIENT_ID,pcse1_0.CLIENT_STORAGE_PROVIDER,pcse1_0.EXTERNAL_CLIENT_ID,pcse1_0.OFFLINE_FLAG,pcse1_0.USER_SESSION_ID,pcse1_0.DATA,pcse1_0.REALM_ID,pcse1_0.TIMESTAMP,pcse1_0.VERSION from OFFLINE_CLIENT_SESSION pcse1_0 where (pcse1_0.CLIENT_ID,pcse1_0.CLIENT_STORAGE_PROVIDER,pcse1_0.EXTERNAL_CLIENT_ID,pcse1_0.OFFLINE_FLAG,pcse1_0.USER_SESSION_ID) in (($1,$2,$3,$4,$5))
select pcse1_0.VERSION from OFFLINE_CLIENT_SESSION pcse1_0 where pcse1_0.USER_SESSION_ID=$1 and pcse1_0.OFFLINE_FLAG=$2 and pcse1_0.CLIENT_ID=$3 and pcse1_0.EXTERNAL_CLIENT_ID=$4 and pcse1_0.CLIENT_STORAGE_PROVIDER=$5 for no key update of pcse1_0 skip locked
update OFFLINE_CLIENT_SESSION set TIMESTAMP=$1,VERSION=$2 where CLIENT_ID=$3 and CLIENT_STORAGE_PROVIDER=$4 and EXTERNAL_CLIENT_ID=$5 and OFFLINE_FLAG=$6 and USER_SESSION_ID=$7 and VERSION=$8
update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1,VERSION=$2 where OFFLINE_FLAG=$3 and USER_SESSION_ID=$4 and VERSION=$5
SET LOCAL synchronous_commit TO OFF
COMMIT
DELETE from JGROUPS_PING WHERE address=$1
INSERT INTO JGROUPS_PING (address, name, cluster_name, ip, coord, last_update, coordinated_by) values ($1, $2, $3, $4, $5, $6, $7)
COMMIT
DELETE from JGROUPS_PING WHERE address=$1
=== 그 sid 를 언급한 SQL — 누가 보냈는가 ===
찾는 sid: jSt9GEPVQLJsO-1CeJjVgltg
2026-09-04 01:12:32.851 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:32.852 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:32.860 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:32.862 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:32.863 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = NULL, $2 = '1788484352', $3 = '{"ipAddress":"10.42.1.50","authMethod":"openid-connect","rememberMe":false,"started":0,"notes":{"KC_DEVICE_NOTE":"
2026-09-04 01:12:32.864 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '{"authMethod":"openid-connect","notes":{"clientId":"131a9912-b578-4b9c-b16a-97518704077e","userSessionStartedAt":"1788484352","iss":"https://aut
2026-09-04 01:12:34.934 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:34.936 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = 'jSt9GEPVQLJsO-1CeJjVgltg', $2 = '0'
2026-09-04 01:12:34.937 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
2026-09-04 01:12:34.938 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = 'jSt9GEPVQLJsO-1CeJjVgltg', $2 = '0', $3 = '131a9912-b578-4b9c-b16a-97518704077e', $4 = 'local', $5 = 'local'
2026-09-04 01:12:34.944 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '1788484354', $2 = '1', $3 = '131a9912-b578-4b9c-b16a-97518704077e', $4 = 'local', $5 = 'local', $6 = '0', $7 = 'jSt9GEPVQLJsO-1CeJjVgltg', $8 =
2026-09-04 01:12:34.946 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '1788484354', $2 = '1', $3 = '0', $4 = 'jSt9GEPVQLJsO-1CeJjVgltg', $5 = '0'
=== 요약: 파드별 질의 건수 ===
6 [keycloak-1]
6 [keycloak-0]
@@ -0,0 +1,26 @@
# 실험 0 — 세션 복제 증거
수집: 2026-09-04 09:54 KST · Keycloak 26 / Infinispan 16.0.12 / PostgreSQL 16
해설: [`docs/experiment-00-session-replication.md`](../../experiment-00-session-replication.md)
| 파일 | 무엇을 보여주는가 |
|---|---|
| `01-cross-node-session.txt` | 클러스터 2멤버 확인 → keycloak-0 로그인 → 같은 sid 가 양쪽에서 보임 → **keycloak-1 이 refresh 성공(200)** → keycloak-1 로그아웃 → **keycloak-0 갱신 실패(400)** → DB 행 삭제 확인 |
| `02-cache-delta.txt` | 로그인 하나를 사이에 둔 양쪽 노드의 캐시 계수기. **keycloak-1 은 전부 +0** |
| `03-cache-ownership.txt` | 로그인을 반대편에 몰아준 결과. **요청을 받은 노드에서만 엔트리가 는다.** 캐시 합 7+5 = DB 12 |
| `session-cache-entries-per-pod.png` | 위 사실의 시계열. 파란 선(keycloak-1)이 0에 붙어 있는 동안 초록 선(keycloak-0)만 14까지 오른다 |
| `keycloak-admin-sessions.png` | 관리 콘솔의 Sessions 화면. 브라우저는 nginx→Traefik 을 거쳐 두 파드 중 하나에 닿지만 **어느 파드가 만든 세션이든 전부 보인다** |
## 핵심 한 줄
클러스터는 형성되지만 **세션 엔트리는 노드를 건너가지 않는다.**
두 노드가 같은 답을 하는 이유는 Infinispan 복제가 아니라 **같은 PostgreSQL** 이다.
| 파일 | 무엇을 보여주는가 |
|---|---|
| `04-read-path-sql.txt` | PostgreSQL 문장 로깅으로 잡은 **keycloak-1 이 실제로 날린 SQL**. `SELECT ... FROM OFFLINE_USER_SESSION` 로 남의 세션을 읽고 `UPDATE ... where VERSION=$5` 로 쓴다. 같은 트랜잭션에 `SET LOCAL synchronous_commit TO OFF` 가 들어 있다 |
## 추론이 관측이 된 지점
0b·0c 는 "keycloak-1 메모리에 없는데 쓸 수 있으니 DB 에서 읽었을 것"이라는
**추론**이었다. 0d 에서 그 SQL 을 파드 IP 와 함께 직접 잡았다.
Binary file not shown.

After

Width:  |  Height:  |  Size: 116 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

+636
View File
@@ -0,0 +1,636 @@
# 실험 0 — 한 노드에서 만든 세션이 다른 노드에서 쓰이는가
로드맵 A-0. 이후 모든 장애 실험의 기준선이다.
> **맥락이 안 잡히면 먼저 읽을 것** —
> [`docs/session-lab-prerequisites.md`](session-lab-prerequisites.md).
> 왜 세션이 문제가 되는지, Keycloak이 세션을 어디에 두는지, 그래서 이 실험이
> 무엇을 가르려는 것인지를 바닥부터 세워둔 문서다.
- 실행 스크립트 — [`deploy/lab/scripts/experiment-session-replication.sh`](../deploy/lab/scripts/experiment-session-replication.sh),
[`experiment-cache-replication-delta.sh`](../deploy/lab/scripts/experiment-cache-replication-delta.sh),
[`experiment-cache-ownership.sh`](../deploy/lab/scripts/experiment-cache-ownership.sh),
[`experiment-session-read-path.sh`](../deploy/lab/scripts/experiment-session-read-path.sh)
- 증거 — [`docs/evidence/session-replication/`](evidence/session-replication/)
- 수집 시각 — 2026-09-04 09:54 KST, Keycloak 26 / Infinispan 16.0.12 / PostgreSQL 16
---
## 0. 결론부터
| 물음 | 답 |
|---|---|
| 한 노드에서 만든 세션을 다른 노드가 쓸 수 있는가 | **그렇다** |
| 로그아웃이 반대 방향으로 전파되는가 | **그렇다** |
| **그 공유는 Infinispan 복제 덕분인가** | **아니다** |
| 그럼 무엇이 공유하는가 | **PostgreSQL** — 반대편 노드가 날린 SQL을 직접 잡았다 |
**클러스터가 형성됐다는 것과 세션이 복제된다는 것은 다른 얘기였다.**
로그에는 `(2) [keycloak-0, keycloak-1]`이 찍히고 `JGROUPS_PING`에도 둘 다
등록되어 있지만, **세션 엔트리는 노드 사이를 건너가지 않는다.**
각 노드는 **자기가 처리한 로그인만** 캐시한다. 두 노드가 같은 답을 내놓는
이유는 복제가 아니라 **같은 데이터베이스를 보기 때문**이다.
---
## 1. 왜 이 실험이 첫 번째인가
앞선 작업에서 Keycloak 2노드 클러스터를 세우고 `ISPN000094`로 멤버 2개를
확인했다. 거기서 멈추면 **"클러스터가 떴다"까지만 아는 것**이고, 그 위에서
장애를 주입해봐야 무엇이 무엇 때문에 깨졌는지 해석할 수 없다.
기준선이 없으면 이런 잘못된 추론을 하게 된다.
> 7800을 막았더니 세션이 깨졌다 → 역시 세션은 7800으로 복제되는구나
실제로는 7800으로 세션이 오가지 않는다는 것을 **먼저** 알아야, 7800을 막았을
때 깨지는 것이 무엇인지 정확히 말할 수 있다.
---
## 2. 실험 설계에서 배운 것 세 가지
측정값보다 **어떻게 측정할지**에서 더 많이 틀렸다. 세 번 고쳤다.
### 2-1. 대조군 없는 측정은 해석할 수 없다
첫 판본은 이렇게 보고했다.
```
=== [4] keycloak-0 이 발급한 토큰을 keycloak-1 이 받는가 ===
http_code=403
```
**403을 "복제 실패"로 읽을 뻔했다.** 발급 노드에도 같은 요청을 보내보니
```
--- userinfo, scope 없음 ---
k0(발급노드) 403
k1(반대편) 403
--- 403 본문 ---
WWW-Authenticate: Bearer realm="master", error="insufficient_scope",
error_description="Missing openid scope"
```
**양쪽 다 403이었다.** 원인은 복제가 아니라 요청에 `openid` scope가 없다는
것이었다. 오히려 **두 노드가 똑같이 답했다는 사실 자체가 일치의 증거**였다.
> **원칙** — 반대편 노드의 응답은 발급 노드의 응답과 나란히 놓기 전까지
> 아무 의미가 없다. 시험군만 재는 측정은 측정이 아니다.
### 2-2. 개수가 아니라 식별자로 추적한다
`client-session-stats``active=2`를 돌려줬다. 그런데 스크립트 자체가
로그인을 두 번 하고(시험용 + 관리 API 호출용) 있었다. **개수는 실험 도구가
만든 잡음에 그대로 오염된다.**
바꾼 방식: 토큰의 `sid`를 뽑아, 각 노드의 세션 목록에 **그 sid가 있는지**를
본다. 개수가 몇이든 상관없다.
```
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
```
### 2-3. 세션 저장소를 실제로 건드리는 탐침을 골라야 한다
| 탐침 | 하는 일 | 적합한가 |
|---|---|---|
| `userinfo` | 서명 검증 + scope 확인 | **아니다.** 세션을 몰라도 통과할 수 있다 |
| **`refresh_token` 그랜트** | 세션을 찾고, 살아있는지 보고, 갱신 시각을 쓴다 | **그렇다** |
refresh는 **읽고 쓴다.** 그래서 "저 노드가 이 세션을 정말로 아는가"에 답한다.
여기에 더해 refresh token은 **회전(rotation)** 된다 — 한 번 쓰면 옛 것이
무효가 된다. 따라서 **반대편 노드에 먼저 써야** 한다. 발급 노드에 먼저 쓰면
시험군에 쓸 토큰이 사라진다. 대조군과 시험군의 순서가 강제된다.
---
## 3. 실험 0 — 교차 노드 세션 사용
### 실행
```bash
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "delete from offline_user_session"
kubectl -n keycloak-lab rollout restart statefulset/keycloak # 캐시를 비운다
./deploy/lab/scripts/experiment-session-replication.sh
```
nginx나 Traefik을 거치지 않고 **파드 IP로 직접** 말을 건다. 로드밸런서를
거치면 어느 노드가 처리했는지가 감춰지는데, 그게 바로 이 실험의 질문이다.
### 결과 — [`01-cross-node-session.txt`](evidence/session-replication/01-cross-node-session.txt)
```
### 사전 확인: 클러스터가 2 멤버로 형성되었는가
ISPN000094: Received new cluster view for channel ISPN:
[keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749, keycloak-0-30843]
name | ip | coord
------------------+-----------------+-------
keycloak-1-48749 | 10.42.0.35:7800 | t
keycloak-0-30843 | 10.42.1.43:7800 | f
=== 대상 ===
keycloak-0 10.42.1.43 kc-lab-2
keycloak-1 10.42.0.35 kc-lab-1
=== [1] keycloak-0 에서 로그인 ===
sid jiv3rVZi1VeaO07oVJkL_MYW
iss https://auth.hyeonworks.com/realms/master
access 수명 60초
refresh 수명 1800초 typ=Refresh
=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
=== [5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용 ===
HTTP 200 ← 기대대로
새 토큰의 sid → 동일 ✔
=== [6] keycloak-1 을 통해 로그아웃 ===
http_code=204
=== [7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파) ===
HTTP 400 ← 기대대로
error invalid_grant
error_description Session not active
=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ===
행 없음 — 로그아웃으로 삭제되었다
```
**네 가지가 모두 기대대로다.**
| | 확인된 것 |
|---|---|
| 조회 | 같은 sid가 양쪽에서 보인다 |
| **쓰기** | keycloak-0의 refresh token을 keycloak-1이 받아 갱신했고, **sid가 유지된다** |
| **역방향 무효화** | keycloak-1의 로그아웃이 keycloak-0의 갱신을 막았다 |
| 영속 | 로그아웃과 함께 DB 행이 사라졌다 |
**`sid`는 JWT 안에만 있는 값이 아니다.** PostgreSQL의
`OFFLINE_USER_SESSION.user_session_id` 컬럼에 **문자 그대로** 들어 있다.
---
## 4. 실험 0b — 복제인가, 같은 DB를 본 것인가
실험 0은 "두 노드가 같은 답을 한다"까지만 증명한다. **그것으로는 Infinispan이
복제했다고 말할 수 없다.** `persistent-user-sessions`(Keycloak 26 기본값)에서는
세션이 PostgreSQL에 기록되므로, **캐시를 아예 꺼도 두 노드는 같은 답을 한다.**
가르는 방법: 로그인 한 번을 사이에 두고 **양쪽 노드의 캐시 계수기**를 잰다.
### 결과 — [`02-cache-delta.txt`](evidence/session-replication/02-cache-delta.txt)
```
=== 로그인은 keycloak-0 에만 보냈다 ===
로그인 응답: http_code=200
=== keycloak-0 (로그인을 받은 노드) ===
계수기 캐시 전 후 증가
approximate_entries_unique sessions 1 2 +1 ←
stores sessions 2 3 +1 ←
misses sessions 3 4 +1 ←
rpc.replication_count sessions 1 1 +0
=== keycloak-1 (아무 요청도 받지 않은 노드) ===
approximate_entries_unique sessions 0 0 +0
stores sessions 1 1 +0
hits sessions 4 4 +0
rpc.replication_count sessions 7 7 +0
```
**keycloak-1의 계수기가 하나도 움직이지 않았다.** 엔트리도 0, 저장도 0.
그리고 keycloak-1의 `sessions` 캐시 엔트리는 **처음부터 끝까지 0**이다.
keycloak-0이 세션을 9개 들고 있는 동안에도 0이었다.
---
## 5. 실험 0c — 엔트리는 어느 노드에 있는가
0b의 결과에는 두 가지 설명이 가능하다.
| | |
|---|---|
| (a) **분산 캐시 + owners=1** | 일관 해싱으로 흩어지는데 이번 건이 우연히 keycloak-0에 떨어졌다 |
| (b) **로컬 캐시** | 각 노드는 자기가 처리한 것만 캐시한다 |
**반대편 노드에 로그인을 몰아주면 갈린다.** (a)라면 어느 쪽에 요청하든 엔트리는
양쪽에 흩어진다. (b)라면 **요청을 받은 노드에서만** 는다.
### 결과 — [`03-cache-ownership.txt`](evidence/session-replication/03-cache-ownership.txt)
```
단계 k0 entries k1 entries
시작 2.0 0.0
keycloak-1 에 로그인 5회 2.0 5.0 ← k0 그대로, k1 만 +5
keycloak-0 에 로그인 5회 7.0 5.0 ← k0 만 +5, k1 그대로
=== 대조: PostgreSQL 에는 몇 건인가 ===
online 세션 12 ← 7 + 5 = 12, 정확히 일치
```
**(b)다.** 그리고 **7 + 5 = 12**로 DB 총계와 정확히 맞는다 — 모든 세션이 DB에
있고, 각각은 **자기를 만든 노드 한 곳에만** 캐시되어 있다.
### 그래프로 본 같은 사실
![세션 캐시 엔트리 수](evidence/session-replication/session-cache-entries-per-pod.png)
`vendor_statistics_approximate_entries_unique{cache="sessions"}` — Grafana Explore.
**파란 선(keycloak-1)이 0에 붙어 있는 동안 초록 선(keycloak-0)만 14까지
올라간다.** 파란 선은 09:50, 즉 **keycloak-1에 직접 로그인을 보낸 순간에만**
5로 뛴다. 중간의 절벽은 캐시를 비우려고 파드를 재시작한 지점이다.
> 캐시 설정은 파일에서 읽을 수 없다. 파드의 `/opt/keycloak/conf/cache-ispn.xml`은
> `<cache-container name="keycloak"><transport/></cache-container>` 뿐이고,
> Keycloak 26은 캐시를 **코드에서** 만든다. 그래서 위 결론은 설정을 읽어서가
> 아니라 **동작을 측정해서** 얻었다.
---
## 6. 실험 0d — 반대편 노드가 정말 DB에서 읽는가
0b·0c까지는 **추론**이었다. "keycloak-1의 메모리에 없는데 쓸 수 있으니 DB에서
읽었을 것이다" — 그럴듯하지만 **SQL을 본 적은 없다.**
PostgreSQL의 문장 로깅을 몇 초만 켜고, keycloak-0에서 만든 세션에 대해
**keycloak-1에 refresh를 딱 한 번** 보낸 뒤 로그를 뒤졌다.
```bash
alter system set log_statement='all';
alter system set log_line_prefix='%m [%p] %h '; -- %h 로 파드 IP 를 남긴다
select pg_reload_conf();
```
### 잡힌 트랜잭션 — [`04-read-path-sql.txt`](evidence/session-replication/04-read-path-sql.txt)
```
01:12:34.934 pid=81376 | BEGIN
01:12:34.934 pid=81376 | select ... from OFFLINE_USER_SESSION where (OFFLINE_FLAG,USER_SESSION_ID) in (($1,$2))
01:12:34.936 pid=81376 | select VERSION from OFFLINE_USER_SESSION ... for no key update skip locked
01:12:34.937 pid=81376 | select ... from OFFLINE_CLIENT_SESSION where (...) in ((...))
01:12:34.938 pid=81376 | select VERSION from OFFLINE_CLIENT_SESSION ... for no key update skip locked
01:12:34.944 pid=81376 | update OFFLINE_CLIENT_SESSION set TIMESTAMP=$1,VERSION=$2 where ... and VERSION=$8
01:12:34.946 pid=81376 | update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1,VERSION=$2 where ... and VERSION=$5
01:12:34.946 pid=81376 | SET LOCAL synchronous_commit TO OFF
01:12:34.947 pid=81376 | COMMIT
```
이 연결의 클라이언트 IP는 `10.42.0.35` — **keycloak-1의 파드 IP**다.
sid 하나에 대해 keycloak-0이 6건(로그인), keycloak-1이 6건(갱신)을 날렸다.
```
=== 요약: 파드별 질의 건수 ===
6 [keycloak-1]
6 [keycloak-0]
```
**추론이 관측이 되었다.** keycloak-1은 세션을 DB에서 읽고, DB에 쓴다.
### 여기서 딸려 나온 것 세 가지
이 13밀리초짜리 트랜잭션 하나에 **원래 질문들의 답이 절반쯤 들어 있다.**
#### (1) 낙관적 락 — `VERSION` 컬럼
```sql
update OFFLINE_USER_SESSION
set LAST_SESSION_REFRESH=$1, VERSION=$2
where OFFLINE_FLAG=$3 and USER_SESSION_ID=$4 and VERSION=$5
```
읽은 뒤 다른 노드가 먼저 고쳤다면 `VERSION`이 달라져 **`UPDATE`가 0행을
갱신하고 실패한다.** 잠금을 오래 잡지 않고 충돌을 사후에 검출하는 방식이다.
**리프레시 토큰 동시 갱신 경쟁(로드맵 B-5)이 여기서 갈린다.** 두 요청이
같은 세션을 동시에 갱신하면 하나는 이 검사에서 진다.
#### (2) `FOR NO KEY UPDATE ... SKIP LOCKED`
```sql
select VERSION from OFFLINE_USER_SESSION
where USER_SESSION_ID=$1 and OFFLINE_FLAG=$2
for no key update of puse1_0 skip locked
( )
```
| 절 | 뜻 |
|---|---|
| `FOR NO KEY UPDATE` | 행을 잠그되 **외래키 참조는 막지 않는다.** `FOR UPDATE`보다 약해 경합이 준다 |
| **`SKIP LOCKED`** | 이미 잠긴 행을 **기다리지 않고 건너뛴다** |
`SKIP LOCKED`가 핵심이다. 같은 세션에 동시 요청이 몰려도 **줄을 서지 않는다.**
대기 대신 낙관적 락 실패로 처리한다 — 처리량을 위해 **지연 대신 재시도**를
고른 설계다.
#### (3) `SET LOCAL synchronous_commit TO OFF` — 내구성을 일부 포기한다
**같은 트랜잭션 안에서**, `COMMIT` 직전에 나온다. pid로 경계를 확인했다.
| | |
|---|---|
| 기본값 `on` | `COMMIT`**WAL이 디스크에 내려간 뒤** 돌아온다 |
| **`off`** | **WAL 플러시를 기다리지 않고** 즉시 돌아온다 |
**결과: PostgreSQL이 갑자기 죽으면 직전 수백 밀리초의 세션 갱신이 사라질 수
있다.** 커밋했다고 응답해놓고 없어진다.
Keycloak이 이걸 의도적으로 켠 이유는 명확하다 — `LAST_SESSION_REFRESH` 갱신은
**초당 수백 번 일어나고, 몇백 밀리초쯤 잃어도 사용자가 다시 갱신하면 그만**이다.
로그인·로그아웃 같은 것과 달리 잃어도 되는 쓰기다.
> **DB 복구 실험(A-2)에서 그대로 관측될 지점이다.** PostgreSQL을 정상 종료가
> 아니라 강제 종료시키면, 마지막 몇백 밀리초의 세션 갱신이 실제로 없어져야
> 한다. 이건 버그가 아니라 **설계된 트레이드오프**다.
```bash
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "show synchronous_commit" # 전역 기본값은 on
```
전역 설정은 `on`이고, **Keycloak이 세션 트랜잭션에만 `SET LOCAL`로 끈다.**
`SET LOCAL`은 그 트랜잭션이 끝나면 되돌아간다.
### 덤: 캐시는 읽어도 채워지지 않는다
```
K1_ENTRIES_BEFORE=5.0
REFRESH_ON_K1=200
K1_ENTRIES_AFTER=5.0 ← 갱신을 처리하고도 그대로
```
**keycloak-1은 남의 세션을 DB에서 읽어 처리하고도 캐시에 담지 않았다.**
0c에서 세운 모델 "각 노드는 자기가 처리한 것만 캐시한다"를 더 좁혀야 한다.
> 캐시에 담기는 것은 **그 노드가 로그인시켜 만든 세션**뿐이다.
> 남의 세션은 매번 DB에서 읽는다.
로드밸런서가 세션을 만든 노드가 아닌 쪽으로 요청을 보내면 **매번 DB를 친다.**
세션 어피니티(sticky session)가 정확성이 아니라 **성능** 문제인 이유가 이것이다.
### 덤 2: jdbc-ping 하트비트가 그대로 보인다
```
01:12:37.551 pid=81369 | BEGIN
01:12:37.551 pid=81369 | DELETE from JGROUPS_PING WHERE address=$1
01:12:37.552 pid=81369 | INSERT INTO JGROUPS_PING (address, name, cluster_name, ip, coord, last_update, coordinated_by) values (...)
01:12:37.553 pid=81369 | COMMIT
```
**디스커버리는 별도 연결(pid=81369)에서 주기적으로 자기 행을 지우고 다시
넣는다.** 세션 트래픽과 완전히 분리된 경로다 — 11층에서 말한 "디스커버리와
트랜스포트는 다른 경로"가 로그에서 눈으로 확인된다.
---
## 7. 그래서 무엇이 세션을 공유하는가
```
로그인 (keycloak-0)
├──▶ PostgreSQL OFFLINE_USER_SESSION ← 진실의 원천. 양쪽이 본다
└──▶ keycloak-0 로컬 캐시 ← 자기 것만. 건너가지 않는다
keycloak-1 이 그 세션을 물으면
└──▶ 자기 캐시에 없음 → PostgreSQL 에서 읽는다
```
| 계층 | 역할 | 노드 간 공유 |
|---|---|---|
| **PostgreSQL** | 진실의 원천 | **여기서 일어난다** |
| **Infinispan `sessions`** | 자기 노드가 처리한 세션의 룩어사이드 캐시 | **일어나지 않는다** |
| **Infinispan 클러스터** | 무효화 메시지, `work` 캐시 등 | 형성은 되어 있다 |
이건 **Keycloak 26의 의도된 설계**다. `persistent-user-sessions`가 기본이 되면서
DB가 진실의 원천이 됐고, 세션 캐시는 **복제할 이유가 없어졌다.** 복제를 하면
네트워크와 메모리를 쓰면서 DB와 캐시 두 벌을 정합하게 유지해야 한다.
---
## 8. 개념
### 8-1. `persistent-user-sessions`
Keycloak 25에서 도입되고 **26에서 기본값**이 된 기능. 사용자 세션을
Infinispan에만 두지 않고 **데이터베이스에 기록**한다.
| | 켜져 있을 때 (기본) | 꺼져 있을 때 (volatile) |
|---|---|---|
| 진실의 원천 | **PostgreSQL** | Infinispan |
| 전체 재시작 후 | **세션이 남는다** | 전부 사라진다 |
| 노드 간 공유 | DB가 한다 | **복제가 해야 한다** |
| 로그인당 비용 | DB 쓰기 | 네트워크 복제 |
**이 실험의 결론은 전부 "켜져 있을 때"의 이야기다.** 끄면 다른 그림이 나오고,
그 비교가 로드맵 A-2다.
```bash
kubectl -n keycloak-lab exec keycloak-0 -- \
/opt/keycloak/bin/kc.sh show-config 2>/dev/null | grep -i feature
```
### 8-2. 온라인 세션이 `OFFLINE_` 테이블에 들어간다
**`USER_SESSION` 테이블은 존재하지 않는다.** 처음에 이걸 찾다가 없어서 당황했다.
```
public | auth_session | table | keycloak
public | jgroups_ping | table | keycloak
public | offline_client_session | table | keycloak
public | offline_user_session | table | keycloak
public | revoked_token | table | keycloak
public | root_auth_session | table | keycloak
```
`persistent-user-sessions`는 **기존 오프라인 세션 테이블을 재사용**하고
`offline_flag` 컬럼으로 구분한다.
| `offline_flag` | 의미 |
|---|---|
| **`'0'`** | **온라인 세션** (일반 로그인) |
| `'1'` | 오프라인 세션 (`offline_access`) |
기본키가 `(user_session_id, offline_flag)` 복합키인 이유다 — 같은 세션 id가
온라인/오프라인 두 행으로 존재할 수 있다.
```sql
select offline_flag, count(*) from offline_user_session group by offline_flag;
select user_session_id, offline_flag, created_on, last_session_refresh
from offline_user_session where user_session_id = '<sid>';
```
**이름이 내용을 배신하는 스키마다.** 운영에서 "온라인 세션이 DB 어디 있냐"를
찾을 때 이걸 모르면 한참 헤맨다.
### 8-3. `sid` — 토큰과 DB를 잇는 열쇠
```
JWT access_token 의 sid jiv3rVZi1VeaO07oVJkL_MYW
↕ 같은 값
DB user_session_id jiv3rVZi1VeaO07oVJkL_MYW
↕ 같은 값
Admin API 세션 목록의 id jiv3rVZi1VeaO07oVJkL_MYW
```
세 곳에서 같은 문자열이다. **장애를 추적할 때 이 값 하나로 토큰·DB·관리 API를
꿰뚫을 수 있다.** 백채널 로그아웃의 `sid` 클레임도 이것이다.
### 8-4. `openid` scope가 없으면 OIDC 토큰이 아니다
`admin-cli``scope` 없이 direct grant를 하면 나오는 클레임은 이렇다.
```
--- access_token ---
클레임: azp, exp, iat, iss, jti, scope, sid, typ
typ = Bearer | sub = None | sid = OxikTqdHCJ7ESm2GPKI1oa6c
```
**`sub`이 없다.** OIDC가 아니라 순수 OAuth2 액세스 토큰이기 때문이다.
`sub`은 OIDC가 요구하는 클레임이고, `openid` scope가 있어야 붙는다.
같은 이유로 `userinfo`가 403 `insufficient_scope`를 준다 — userinfo는 OIDC
엔드포인트다. **두 현상은 하나의 원인**이다.
### 8-5. 룩어사이드(lookaside) 캐시
```
읽기: 캐시 확인 → 없으면 DB → 캐시에 채움
쓰기: DB 에 쓰고 → 캐시에도 씀
```
캐시가 **DB 앞에 서 있되 DB를 대체하지 않는** 구조. 캐시를 통째로 날려도
정확성은 유지되고 느려지기만 한다. Keycloak 26의 세션 캐시가 이 모양이다.
이 성질이 **노드 상실 실험(A-3)의 결과를 미리 결정한다** — 노드가 죽으면
그 노드의 캐시는 사라지지만 세션은 DB에 있으므로 살아남아야 한다.
---
## 9. 다음 실험에 대한 예측
기준선이 생겼으므로 **틀릴 수 있는 예측**을 세울 수 있다. 예측이 빗나가면
그것이야말로 배울 거리다.
| 실험 | 예측 | 근거 |
|---|---|---|
| **A-1** TCP 7800 차단 | **세션 공유는 안 깨진다.** 대신 무효화 전파와 `work` 캐시가 깨진다 | 세션은 7800으로 오가지 않는다 |
| **A-2** DB 손실 | **즉시 전면 장애.** 캐시에 있는 세션도 못 쓴다 | DB가 진실의 원천 |
| **A-2'** DB **강제** 종료 | 직전 수백 ms 의 세션 갱신이 **사라진다** | `synchronous_commit OFF` |
| **B-5** 동시 갱신 경쟁 | 한쪽이 `VERSION` 검사에서 지고 재시도한다 | 낙관적 락 |
| **A-3** 노드 상실 (kc-lab-2) | **세션은 살아남는다.** 죽은 노드의 캐시만 사라진다 | 룩어사이드 |
| **A-4** volatile 비교 | 7800 차단이 **A-1과 정반대로** 치명적이 된다 | 그때는 캐시가 진실의 원천 |
특히 A-1은 **직관과 어긋나는 예측**이다. "클러스터 포트를 막으면 세션이
깨진다"가 상식이지만, 이 기준선이 맞다면 안 깨져야 한다.
---
## 10. 겪은 함정
### 10-1. kubectl 스트림에서 출력이 통째로 사라졌다
`kubectl run --rm -i ... | grep` 로 받으면 **중간 조각이 유실됐다.**
keycloak-1의 스냅샷과 그 다음 마커가 함께 없어져, 전값이 0으로 잡히면서
**가짜 델타가 만들어졌다.**
```
###BEFORE_K1 ← 여기 있어야 할 지표 20줄과
http_code=200 다음 마커 ###LOGIN 이 통째로 사라졌다
###AFTER_K0
```
이때 리포트는 keycloak-1이 `+9`, `+7` 증가한 것처럼 보였다. **없는 복제가
있는 것처럼 보이는, 가장 나쁜 종류의 오류다.**
| 고친 방법 | |
|---|---|
| 파드 안에서 파일로 모으고 마지막에 `cat` 한 번 | 스트리밍 중 유실을 없앤다 |
| 스냅샷이 비면 **경고를 출력**한다 | 조용히 0으로 계산되는 것을 막는다 |
```bash
for n in ('BEFORE_K0','BEFORE_K1','AFTER_K0','AFTER_K1'):
if not blocks.get(n):
print(f' !! {n} 스냅샷이 비었다 — 델타를 신뢰할 수 없다')
```
**계측 코드는 자기가 실패했는지 스스로 말해야 한다.**
### 10-2. DB에서 직접 지우면 캐시는 남는다
정리하려고 `delete from offline_user_session`을 실행했더니, **캐시 엔트리는
그대로 남아** 캐시 합계(19)와 DB 총계(15)가 어긋났다.
> 운영에서 세션 테이블을 직접 손대면 캐시와 DB가 갈라진다. 세션을 지울 때는
> 관리 API(`logout-all`)를 쓰거나, DB를 건드렸다면 **파드를 재시작**해야 한다.
이 실험의 최종 수치는 **파드 재시작 후** 다시 잰 것이다.
### 10-3. Keycloak 이미지에는 `curl`이 없다
`kubectl exec keycloak-0 -- curl` 은 실패한다. 임시 `curlimages/curl` 파드를
띄워 파드 네트워크 안에서 호출했다. 파드 IP는 클러스터 밖에서 닿지 않으므로
이 방법이 사실상 유일하다.
### 10-4. 중첩 셸의 변수 치환
`ssh host '... $VAR ...'` 안에 다시 `sh -c "..."` 를 넣으면 인용이 세 겹이 되어
치환이 조용히 깨진다. 첫 시도에서 파드 IP가 빈 문자열이 되어 아무 출력도
나오지 않았다.
**스크립트 파일로 만들어 `scp` 로 옮기는 쪽이 옳다.** 재현도 되고 저장소에
남는다. `deploy/lab/scripts/` 아래 세 스크립트가 그 결과다.
---
## 11. 재현
```bash
# 1. 깨끗한 상태로 되돌린다 (DB 비우고 캐시 비우기)
ssh test-server '
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "delete from offline_user_session"
kubectl -n keycloak-lab rollout restart statefulset/keycloak
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=300s'
# 2. 세 실험을 순서대로
ssh test-server '/tmp/experiment-session-replication.sh' # 교차 노드 사용
ssh test-server '/tmp/experiment-cache-replication-delta.sh' # 복제인가 DB인가
ssh test-server '/tmp/experiment-cache-ownership.sh' # 엔트리 위치
# 3. 그래프
# https://app2.hyeonworks.com/explore
# vendor_statistics_approximate_entries_unique{cache="sessions"}
# Legend: {{pod}} on {{node}}
```
### 확인용 명령 모음
```bash
# 클러스터 멤버
kubectl -n keycloak-lab logs keycloak-0 | grep ISPN000094 | tail -1
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "select name, ip, coord from jgroups_ping"
# DB 세션
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "select offline_flag, count(*) from offline_user_session group by offline_flag"
# 노드별 캐시 엔트리 (파드 안에서)
curl -s http://<pod-ip>:9000/metrics \
| grep 'approximate_entries_unique{cache="sessions"'
```
@@ -0,0 +1,506 @@
# A-1 — 노드 간 통신(TCP 7800)을 끊으면 무엇이 깨지는가
브랜치 `feature/keycloak-a1-jgroups-transport-block` ·
증거 [`docs/evidence/a1-jgroups-transport-block/`](evidence/a1-jgroups-transport-block/) ·
2026-09-04 11:3811:52 KST · Keycloak 26.7.0 / Infinispan 16.0.12
맥락은 [`session-lab-prerequisites.md`](session-lab-prerequisites.md),
기준선은 [`experiment-00-session-replication.md`](experiment-00-session-replication.md).
---
## 0. 결론부터
| 예측 | 결과 |
|---|---|
| 세션 공유는 **안 깨진다** | **맞다.** 교차 노드 refresh 가 `200` |
| 로그아웃 전파는 **안 깨진다** | **틀렸다.** `400` 이어야 할 것이 `200` |
| — | **NetworkPolicy 만으로는 분단이 일어나지 않는다** (예상 못 함) |
| — | **분단된 노드가 스스로 로드밸런서에서 빠진다** (예상 못 함) |
**예측 하나가 빗나갔고, 예상하지 못한 것이 둘 나왔다.** 그중 하나는
실험 방법 자체를 무효화할 뻔했다.
---
## 1. 왜 이 실험인가
A-0 에서 **세션은 Infinispan 복제가 아니라 PostgreSQL 로 공유된다**는 것을
측정했다. 그렇다면 통념과 정면으로 어긋난다.
| | |
|---|---|
| **통념** (Keycloak 24 이전 자료) | 세션은 7800 으로 복제된다 → **막으면 세션 공유가 깨진다** |
| **A-0 측정** | 세션은 DB 로 공유된다 → **막아도 안 깨진다** |
둘 중 하나는 틀렸고, 이 실험이 판정한다.
---
## 2. 기준선
```
=== [기준선 1] 클러스터 뷰 ===
keycloak-0: [keycloak-1-48749|5] (2) [keycloak-1-48749, keycloak-0-30843]
keycloak-1: [keycloak-1-48749|5] (2) [keycloak-1-48749, keycloak-0-30843]
=== [기준선 2] JGROUPS_PING ===
keycloak-0-30843 | 10.42.1.43:7800 | f
keycloak-1-48749 | 10.42.0.35:7800 | t ← 코디네이터는 하나
=== [기준선 4] JGroups 지표 (양쪽 동일) ===
fd_sock2_get_num_suspected_members 0.0
merge3_get_num_merge_events 0.0
nakack2_get_xmit_table_missing 0.0
```
**대조군** — 차단 전에 같은 절차를 그대로 한 번 돌린다.
```
=== [대조군] keycloak-0 로그인 → keycloak-1 에서 refresh ===
sid tAWs2gCPr6SOcD4jDR9-_CzB
keycloak-1 에서 refresh: 200
```
A-0 에서 배운 규칙이다 — **시험군만 재는 측정은 측정이 아니다.**
---
## 3. 주입 — NetworkPolicy 로 7800 만 막는다
```bash
kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
```
```yaml
spec:
podSelector: { matchLabels: { app: keycloak } }
policyTypes: [Ingress]
ingress:
- ports:
- { port: 8080, protocol: TCP } # HTTP — 열어둔다
- { port: 9000, protocol: TCP } # health+metrics — 열어둔다
# 7800 은 일부러 없다
```
### 개념 — NetworkPolicy 는 방화벽이 아니라 **허용 목록**이다
**"7800 을 거부"라고 쓸 수 없다.** 파드가 `policyTypes: [Ingress]` 를 가진
정책에 선택되는 순간 **모든 인바운드가 거부**되고, 규칙에 적힌 것만 통과한다.
그래서 7800 은 **빠뜨림으로써** 막힌다.
이 구조가 두 허용 규칙을 **결정적으로 만든다.** 잘못 쓰면 분단된 클러스터가
아니라 **죽은 Keycloak 을 측정하게 된다.**
| 포트 | 빼면 |
|---|---|
| 8080 | Traefik·상대 노드의 REST 호출이 전부 끊긴다 |
| **9000** | **readiness 프로브가 실패해 kubelet 이 파드를 죽인다** — 엉뚱한 이유로 클러스터가 깨진다 |
적용 직후 확인했다.
```
파드 상태: keycloak-0 ready=true restarts=0
keycloak-1 ready=true restarts=0
9000 도달: 10.42.1.43:9000 health=200 / 10.42.0.35:9000 health=200
8080 도달: 10.42.1.43:8080 root=200 / 10.42.0.35:8080 root=200
```
**주입이 의도한 것만 건드렸음을 먼저 확인한 뒤에 결과를 해석한다.**
---
## 4. 문제 ① — **NetworkPolicy 만으로는 분단이 안 된다**
가장 중요한 발견이며, 하마터면 **실험 전체를 무효로 만들 뻔했다.**
차단 후 지표가 꿈쩍도 하지 않았다. 신규 연결은 분명히 막히는데.
```
=== 7800 신규 연결 ===
10.42.1.43:7800 curl exit=7 (연결 실패)
10.42.1.43:9000 curl exit=28 (연결됨, telnet 이라 대기 → 타임아웃)
```
그런데 파드 내부 소켓을 보니
```
=== /proc/net/tcp6 · 7800 = 0x1E78 ===
keycloak-0: ...2B012A0A:1E78 ...23002A0A:9C57 01 ← 01 = ESTABLISHED
keycloak-1: ...23002A0A:9C57 ...2B012A0A:1E78 01
(10.42.0.35:40023 → 10.42.1.43:7800)
```
**기존 연결이 멀쩡히 살아 있다.**
### 왜 그런가 — conntrack
```
패킷 도착
├─▶ [ conntrack: ESTABLISHED/RELATED 이면 ACCEPT ] ← 여기서 통과해버린다
└─▶ [ NetworkPolicy 규칙 평가 ] ← 여기까지 오지 않는다
```
리눅스 방화벽은 성능을 위해 **이미 성립한 연결을 먼저 통과**시킨다.
NetworkPolicy 는 그 뒤에 있으므로 **신규 연결(SYN)만** 걸러낸다.
```
=== conntrack 확인 ===
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 ... [ASSURED]
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 ... [ASSURED]
tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED]
─────────────────────────────────────────────────────────────
신규 연결은 응답을 못 받는다 = 정책이 동작하고는 있다
```
> **운영적 함의 — NetworkPolicy 는 이미 붙어 있는 것을 떼어내지 못한다.**
> 보안 사고 대응으로 "지금 당장 이 통신을 끊어라"에 NetworkPolicy 를 적용하면,
> **새 연결만 막히고 진행 중인 연결은 계속된다.** 끊으려면 conntrack 을 지우거나
> 파드를 재시작해야 한다.
### 덤 — **57800 포트도 있다**
`sport=50477 dport=57800` — FD_SOCK2 는 **`bind_port + 50000`** 을 쓴다.
7800 만 막고 57800 을 열어두면 장애 감지 채널이 남는다.
이 실험의 허용 목록 방식은 **둘 다 자동으로 막았다** — 8080·9000 외 전부 거부이므로.
### 조치
```bash
# 정확한 튜플로 지정해야 지워진다. --dport 만으로는 0건이었다
sudo conntrack -D -p tcp -s 10.42.0.35 -d 10.42.1.43 --sport 40023 --dport 7800
sudo conntrack -D -p tcp -s 10.42.1.43 -d 10.42.0.35 --sport 7800 --dport 40023 # 역방향
```
**양쪽 노드에서, 양쪽 방향으로** 지워야 한다. 서버 쪽 노드에는 튜플이 뒤집혀
기록되어 있다.
그리고 **즉시 끊기지 않는다.**
```
11:41 conntrack 삭제
11:44 cluster_size 2 → 1 ← 약 3분 뒤
```
TCP 는 상대가 사라졌음을 **재전송 타임아웃**으로 알아낸다. 소켓은 한동안
`ESTABLISHED` 로 남아 있다.
---
## 5. 문제 ② — 계측 도구가 잘못됐다
임시 curl 파드로 20초마다 지표를 긁었더니 이런 결과가 나왔다.
```
+20초 suspected(k0 k1) = []
+60초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
+140초 suspected(k0 k1) = [0.0 ]
```
**빈 값, 개수가 맞지 않는 값이 섞인다.** `kubectl run --rm` 은 매번 파드를
만들고 지우므로 느리고 경합이 있다.
게다가 첫 시도의 판정 조건이
```sh
[ "$R" != "0.0 0.0 " ] && echo "→ 변화 감지" && break
```
여서 **빈 문자열을 "변화"로 읽고 즉시 빠져나왔다.** A-0 에서 똑같은 실수를
했는데 또 했다.
> **임시 파드는 계측 도구가 아니다.** 15초마다 이미 긁고 있는 Prometheus 가
> 그러라고 있는 것이다.
```bash
kubectl -n observability port-forward svc/prometheus 19090:9090 &
curl -s "http://localhost:19090/api/v1/query_range?query=vendor_cluster_size&start=$START&end=$END&step=60"
```
그리고 이 과정에서 **`vendor_cluster_size`** 를 발견했다 — 멤버 수를 직접
알려주는 지표다. 처음부터 이걸 봤어야 했다.
```bash
curl -s "http://localhost:19090/api/v1/label/__name__/values" | grep -E "cluster|member|view"
```
---
## 6. 진짜 분단이 일어난 순간
```
=== vendor_cluster_size ===
keycloak-1: 11:43:57=2 11:44:27=1 ... 11:51:28=2
keycloak-0: 11:43:57=2 (파드 교체) 11:45:27=1 ... 11:51:28=2
```
![cluster_size 추이](evidence/a1-jgroups-transport-block/a1-cluster-size-partition-recovery.png)
정책이 걸린 채 `keycloak-0` 이 재시작되자, 로그가 정확히 말해준다.
```
GMS: JOIN(keycloak-0-26403) sent to keycloak-1-48749 timed out ← 10회
GMS: too many JOIN attempts (10): becoming singleton ← 포기
ISPN000094: new cluster view [keycloak-0-26403|0] (1) [keycloak-0-26403]
```
`keycloak-1` 쪽도 혼자가 되었다.
```
ISPN000094: [keycloak-1-48749|6] (1) [keycloak-1-48749]
```
### **DB 에는 둘 다 있는데 클러스터는 안 붙는다** — 예측한 그 상태
```
=== JGROUPS_PING ===
name | ip | coord
------------------+-----------------+-------
keycloak-0-26403 | 10.42.1.67:7800 | t ← 코디네이터
keycloak-1-48749 | 10.42.0.35:7800 | t ← 코디네이터
```
**`coord = t` 가 둘.** 교과서적인 split brain 이며, **데이터베이스 한 줄로
확인된다.** 디스커버리(DB)는 살아 있고 트랜스포트(7800)만 죽은 상태다.
**단일 노드에서는 만들 수 없는 고장**이며, 이 실험대를 2 VM 으로 만든 이유다.
---
## 7. 본 시험 — 분단 상태에서 세션은 어떻게 되는가
```
[1] keycloak-0 로그인 sid=nShl5TaBrZnKStDqaspjgmJB
[2] keycloak-1 에서 refresh HTTP 200 ← 예측대로
[3] keycloak-1 에서 로그아웃 HTTP 204
[4] keycloak-0 에서 재갱신 시도 HTTP 200 ← 400 이어야 했다
```
### [2] 세션 공유 — **예측이 맞았다**
클러스터가 갈라졌는데도 **한쪽에서 만든 세션을 반대쪽이 갱신했다.**
A-0 의 모델이 맞고, **통념이 틀렸다.** 세션은 7800 으로 다니지 않는다.
### [4] 로그아웃 전파 — **예측이 틀렸다**
A-0 에서는 같은 절차가 `400 invalid_grant / Session not active` 였다.
분단 상태에서는 `200` 이다. **로그아웃한 세션이 반대편에서 살아 있다.**
기제를 확정했다.
```
=== 그 sid 가 DB 에 남아 있는가 ===
user_session_id | offline_flag | last_session_refresh
-----------------+--------------+----------------------
(0 rows) ← DB 행은 삭제되었다
=== 노드별 세션 캐시 엔트리 ===
keycloak-1 kc-lab-1 = 0
keycloak-0 kc-lab-2 = 1 ← 캐시에는 남아 있다
```
```
keycloak-1 로그아웃
├──▶ PostgreSQL 행 삭제 ✔ 되었다
└──▶ keycloak-0 에게 "캐시에서 지워라" ✗ 7800 이 막혀 못 갔다
keycloak-0 은 자기 캐시로 200 을 준다 ◀────────────┘
```
### **A-0 의 결론을 정정한다**
A-0 에서 나는 이렇게 썼다.
> 로그아웃과 함께 DB 행이 사라졌다 → 무효화가 DB 삭제로 전파된다
**그 인과는 틀렸다.** DB 행 삭제는 일어나지만, **반대편 노드는 DB 를 다시
읽지 않는다.** 자기 캐시에 있으면 그걸로 답한다.
> **룩어사이드 캐시는 읽을 때 DB 와 대조하지 않는다.**
> 캐시 무효화는 **클러스터 메시지(7800)를 타고** 간다.
A-0 에서 400 이 나온 것은 DB 덕분이 아니라 **그때는 7800 이 살아 있어서**였다.
두 실험을 붙여야 비로소 정확한 그림이 나온다.
| | 세션 **조회** | 세션 **무효화** |
|---|---|---|
| 경로 | PostgreSQL | **클러스터 메시지 (7800)** |
| 7800 차단 시 | 정상 | **전파되지 않음** |
---
## 8. 그런데 안전장치가 있었다 — 예상 못 한 발견
`keycloak-0``Ready=false` 였다. 이유를 물었더니
```json
{ "status": "DOWN",
"checks": [
{ "name": "Keycloak cluster health check", "status": "DOWN",
"data": { "Failing since": "2026-09-04 02:45:14,251" } },
{ "name": "Keycloak database connections async health check", "status": "UP" }
] }
```
**Keycloak 은 클러스터 분단을 readiness 로 신고한다.** 그리고 쿠버네티스가
그 신고를 받아 처리했다.
```
=== Service 엔드포인트 ===
ready 주소: [10.42.0.35] ← keycloak-1 만 트래픽을 받는다
notReady : [10.42.1.67] ← keycloak-0 은 제외되었다
=== 외부 진입점 ===
https://auth.hyeonworks.com/realms/master HTTP 200
토큰 발급 HTTP 200
```
**분단된 노드가 스스로 로드밸런서에서 빠졌고, 서비스는 계속되었다.**
### 그래서 7절의 로그아웃 우회는 어떻게 봐야 하나
| | |
|---|---|
| 내가 한 것 | Service 를 우회해 **파드 IP 로 직접** 호출 |
| 실제 사용자 | nginx → Traefik → **Service** → Ready 인 파드만 |
**정문으로 들어오면 낡은 캐시에 닿지 않는다.** readiness 게이트가 막는다.
> 다만 이건 **비대칭이라서 살았다.** `keycloak-1` 은 원래 뷰에서 멤버가 하나
> 줄어든 정상적인 사건이라 Ready 를 유지했고, `keycloak-0` 은 합류 자체를
> 못 해 DOWN 이 되었다. **양쪽이 동시에 DOWN 이 되는 경로가 있다면 전면 장애다.**
> A-5(비대칭 파티션)에서 이어서 본다.
---
## 9. 복구
```bash
kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
```
```
+30초 keycloak-0=1 keycloak-1=1
+60초 keycloak-0=1 keycloak-1=1
+90초 keycloak-0=2 keycloak-1=2 ← 재형성
```
**90초 만에 자동으로 다시 붙었다. 사람 손이 필요 없었다.**
```
=== MERGE3 가 합쳤는가 ===
merge_events keycloak-0 = 1
merge_events keycloak-1 = 1
```
**MERGE3 가 한 일이다.** split brain 을 감지해 뷰를 병합하는 프로토콜이며,
지표가 `0 → 1` 로 올라간 것이 그 증거다.
```
=== JGROUPS_PING ===
keycloak-0-26403 | 10.42.1.67:7800 | t
keycloak-1-48749 | 10.42.0.35:7800 | f ← 코디네이터가 하나로 돌아왔다
```
**코디네이터가 keycloak-1 에서 keycloak-0 으로 넘어갔다.** 코디네이터는
특권이 아니라 역할이며, 병합 시 재선출된다.
---
## 10. 개념 정리
### conntrack — 연결 추적
리눅스 커널이 **진행 중인 연결을 기억**하는 표. 패킷마다 규칙을 다시 평가하지
않기 위해 존재한다.
| 상태 | 뜻 |
|---|---|
| `NEW` | 첫 패킷(SYN) |
| **`ESTABLISHED`** | **양방향 통신이 성립함 — 규칙 평가를 건너뛴다** |
| `[ASSURED]` | 충분히 오래된 연결. 표가 꽉 차도 안 지워진다 |
| `SYN_SENT [UNREPLIED]` | 보냈는데 답이 없음 = **차단되고 있다** |
```bash
sudo conntrack -L | grep 7800
sudo conntrack -D -p tcp -s <src> -d <dst> --sport <sp> --dport <dp>
```
### FD_SOCK2 와 포트 규약
| 프로토콜 | 포트 | 하는 일 |
|---|---|---|
| TCP (트랜스포트) | **7800** | 클러스터 메시지 |
| **FD_SOCK2** | **57800** = 7800 + 50000 | 소켓으로 상대 생존 감시 |
**방화벽 규칙을 손으로 쓸 때 57800 을 빠뜨리기 쉽다.**
### MERGE3
split brain 이 생긴 뒤 **갈라진 뷰를 다시 합치는** JGroups 프로토콜.
주기적으로 다른 코디네이터의 존재를 확인하고, 발견하면 병합을 개시한다.
```promql
vendor_jgroups_merge3_get_num_merge_events
```
### readiness 프로브와 Service 엔드포인트
```
readiness 실패 → 파드가 Service 의 notReadyAddresses 로 이동
→ kube-proxy 가 그 파드로 라우팅하지 않음
→ 살아 있지만 트래픽은 안 받음
```
**liveness 와 다르다.** liveness 실패는 **재시작**, readiness 실패는
**격리**다. 클러스터 분단처럼 "재시작해도 안 나아지는" 문제에는 readiness 가
맞는 신호다.
---
## 11. 재현 절차 (명령어)
```bash
# 0. 기준선
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "select name, ip, coord from jgroups_ping order by name"
kubectl -n observability port-forward svc/prometheus 19090:9090 &
curl -s "http://localhost:19090/api/v1/query?query=vendor_cluster_size"
# 1. 차단
kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
# 2. 주입이 의도한 것만 건드렸는지 확인 (8080/9000 은 살아 있어야 한다)
kubectl -n keycloak-lab get pods -o wide | grep keycloak # restarts=0 확인
# 3. 기존 연결이 남아 있음을 확인 — 이걸 안 하면 실험이 무효다
ssh kc-lab-1 'sudo conntrack -L | grep 7800'
# 4. conntrack 삭제 (양쪽 노드, 양쪽 방향). 반영까지 약 3분
ssh kc-lab-1 'sudo conntrack -D -p tcp -s <k1ip> -d <k0ip> --sport <sp> --dport 7800'
ssh kc-lab-2 'sudo conntrack -D -p tcp -s <k0ip> -d <k1ip> --sport 7800 --dport <sp>'
# 5. 분단 확인
curl -s "http://localhost:19090/api/v1/query?query=vendor_cluster_size"
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "select name, coord from jgroups_ping" # coord=t 가 둘이면 split brain
# 6. 복구
kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
curl -s "http://localhost:19090/api/v1/query?query=vendor_jgroups_merge3_get_num_merge_events"
```
---
## 12. 다음 실험에 남기는 것
| 실험 | 이 실험이 준 것 |
|---|---|
| **A-5** 비대칭 파티션 | **양쪽이 동시에 NotReady 가 되는 경로가 있는가.** 여기서는 비대칭이라 살았다 |
| **A-2** DB 정지 | 캐시가 DB 와 대조하지 않는다는 사실 → **캐시에 있는 세션은 DB 없이도 읽힐 수 있다** |
| **A-7** volatile 비교 | 같은 주입에서 세션 공유가 **깨져야** 한다. 이 실험이 그 대조군 |
| 전체 | **주입이 실제로 걸렸는지 먼저 확인한다.** NetworkPolicy 는 기존 연결을 못 끊는다 |
+308
View File
@@ -0,0 +1,308 @@
# A-2 — PostgreSQL 이 죽으면 어떻게 되는가
브랜치 `feature/keycloak-a2-database-loss` ·
증거 [`docs/evidence/a2-database-loss/`](evidence/a2-database-loss/) ·
2026-09-04 11:5611:58 KST · Keycloak 26.7.0
선행: [`A-0`](experiment-00-session-replication.md) ·
[`A-1`](experiment-a1-jgroups-transport-block.md)
---
## 0. 결론부터
| 예측 | 결과 |
|---|---|
| 즉시 전면 장애 | **맞다.** 외부 진입점 **503**, 양쪽 노드 NotReady |
| 캐시에 있어도 못 쓴다 | **맞다.** 캐시를 가진 노드도 `500` |
| — | **`up = 1` 인 채로 전면 장애가 났다** (관측의 함정) |
| — | **DB 복귀 15초 만에 재시작 없이 자동 회복** |
**A-1 과 정반대다.** A-1 은 한쪽만 빠지고 서비스가 계속됐지만,
A-2 는 **살아남는 노드가 없다.**
---
## 1. 설계 — 네 경로를 구분해서 본다
A-1 에서 **"룩어사이드 캐시는 읽을 때 DB 와 대조하지 않는다"** 를 확인했다.
그렇다면 캐시를 가진 노드는 DB 없이도 버틸지 모른다. 그 가설을 가른다.
| # | 경로 | 무엇을 보는가 |
|---|---|---|
| ① | **캐시를 가진 노드**에서 refresh | 캐시가 DB 를 대신할 수 있는가 |
| ② | 캐시가 없는 노드에서 refresh | 완전한 DB 의존 |
| ③ | 새 로그인 | 쓰기 경로 |
| ④ | 이미 발급된 토큰으로 조회 | 서명만으로 되는 경로 |
**access token 수명이 60초**이므로, 토큰 발급 → DB 정지 → 시험을 그 안에
끝내야 한다.
### 계측 도구를 바꿨다
A-1 에서 임시 curl 파드가 형편없는 계측 도구임을 확인했다. 여기서는
**상주 탐침 파드**를 하나 띄우고 `exec` 로 단계를 이어간다. 토큰을 파드 안
파일에 남겨 **DB 정지 전후로 같은 토큰**을 쓸 수 있다.
```bash
kubectl -n keycloak-lab run a2-probe --image=curlimages/curl:8.11.1 \
--restart=Never --command -- sleep 7200
kubectl -n keycloak-lab wait --for=condition=Ready pod/a2-probe --timeout=120s
```
---
## 2. 기준선
```
keycloak-0 ready=true 10.42.1.67 kc-lab-2
keycloak-1 ready=true 10.42.0.35 kc-lab-1
postgres ready=true 10.42.1.24 kc-lab-2
cluster_size keycloak-0 = 2
cluster_size keycloak-1 = 2
```
세션을 양쪽에 하나씩 만들고, A-0 대로 **각자 자기 노드에만 캐시**되는 것을
확인했다.
```
keycloak-0 에서 로그인 sid=EAXV5HcG2J1BZ3vnwONf64AQ
keycloak-1 에서 로그인 sid=McyTj5lj3n_JqApCXeuAHExc
→ 캐시 keycloak-0 = 1 건 / keycloak-1 = 0 건 (스크레이프 지연)
```
---
## 3. 주입
```bash
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=90s
```
```
정지 시각: 11:56:04
삭제 완료: 11:56:04 ← 즉시
```
---
## 4. 결과 — 네 경로
```
① 캐시를 가진 노드(keycloak-0)에서 refresh HTTP 500
② 캐시가 없는 노드(keycloak-1)에서 refresh HTTP 500
③ 새 로그인 HTTP 500
④ 관리 API (세션 조회 필요) HTTP 500
--- 오류 본문 ---
{"error":"unknown_error","error_description":"For more on this error consult the server log."}
```
### ① 이 500 인 것이 중요하다
**캐시에 세션을 들고 있어도 refresh 는 실패한다.**
A-1 에서는 로그아웃된 세션을 캐시로 `200` 을 줬다. 왜 여기서는 안 되는가.
```
refresh 처리
├── 세션이 존재하는가 → 캐시로 답할 수 있다
└── LAST_SESSION_REFRESH 갱신 → DB 쓰기가 필요하다 ← 여기서 죽는다
```
A-0 에서 잡은 SQL 그대로다.
```sql
update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1, VERSION=$2 where ...
```
> **캐시는 읽기를 대신할 뿐, 쓰기를 대신하지 못한다.**
> refresh 는 이름과 달리 **쓰기 연산**이다.
### 로그가 말하는 원인
```
Caused by: java.net.ConnectException: Connection refused
at org.postgresql.core.v3.ConnectionFactoryImpl.tryConnect
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call
```
`agroal` 은 Quarkus 의 커넥션 풀이다. 풀이 새 커넥션을 만들지 못한다.
---
## 5. 살아남은 것 — 상태가 필요 없는 경로
```
JWKS 엔드포인트(realm 공개키) HTTP 200
realm 메타데이터(.well-known) HTTP 200
관리 API (세션 조회 필요) HTTP 500
```
**realm 공개키와 메타데이터는 메모리에 있으므로 DB 없이도 응답한다.**
이론적으로는 **이미 JWKS 를 캐시한 리소스 서버는 토큰 검증을 계속할 수 있다**는
뜻이다. 다만 이 실험대에는 독립 리소스 서버가 아직 없으므로 **여기까지가
말할 수 있는 범위**다 — B층에서 확인한다.
> **그런데 정문으로는 이것도 못 쓴다.** 아래 6절 때문이다.
---
## 6. 전면 장애 — 살아남는 노드가 없다
```
=== 파드 Ready ===
keycloak-0 false restarts=0
keycloak-1 false restarts=0
=== Service 엔드포인트 ===
ready : [] ← 비었다
notReady: [10.42.0.35 10.42.1.67]
=== 외부 진입점 ===
https://auth.hyeonworks.com/realms/master HTTP 503
```
```json
{ "status": "DOWN",
"checks": [
{ "name": "Keycloak cluster health check", "status": "UP" },
{ "name": "Keycloak database connections async health check", "status": "DOWN" },
{ "name": "Keycloak Initialized", "status": "UP" } ] }
```
**`cluster health` 는 UP 인데 `database connections` 가 DOWN 이라 전체가 DOWN 이다.**
헬스체크는 **모든 항목이 UP 이어야 UP** 이다.
### A-1 과의 대비가 이 실험의 핵심이다
| | A-1 (7800 차단) | **A-2 (DB 정지)** |
|---|---|---|
| Ready 인 파드 | keycloak-1 **1개 생존** | **0개** |
| Service `ready` | `[10.42.0.35]` | **`[]`** |
| 외부 응답 | **200** | **503** |
| 성격 | 용량 저하 | **전면 장애** |
**노드를 몇 대로 늘려도 DB 가 죽으면 전부 같이 죽는다.**
Keycloak 의 대수는 DB 장애에 아무 도움이 되지 않는다.
> 원래 질문 *"Redis 또는 DB가 뒤질 경우 어떻게 복구를 해야 되는지"* 에 대한
> 첫 번째 답 — **복구 이전에, DB 이중화가 Keycloak 대수보다 우선한다.**
---
## 7. 관측의 함정 — `up = 1` 인 채로 전면 장애
```
up{pod=keycloak-1} = 1
up{pod=keycloak-0} = 1 ← 서비스는 503 인데
```
![up 은 움직이지 않았다](evidence/a2-database-loss/a2-up-stayed-1-during-outage.png)
**전 구간 평평하다.** (11:44 의 짧은 골은 A-1 에서 파드를 교체한 자국이다.)
`up` 은 **Prometheus 가 `/metrics` 를 긁는 데 성공했는가**만 말한다.
프로세스는 멀쩡히 살아 메트릭을 내놓고 있었다. **기능은 전멸했는데.**
| 지표 | 이 장애에서 |
|---|---|
| `up` | **1 — 아무것도 알려주지 않는다** |
| 파드 `Ready` | **false — 여기서 드러난다** |
| 외부 HTTP 코드 | **503 — 사용자가 겪는 것** |
> **A-0 에서 나는 `up` 을 "가장 중요한 합성 지표"라고 썼다.**
> 절반만 맞다. `up` 은 **대상이 사라진 것**을 잡지만 **대상이 살아서 못 쓰는 것**은
> 못 잡는다. 후자가 운영에서 훨씬 흔하다.
>
> **알림은 `up` 이 아니라 readiness 와 외부 응답 코드에 걸어야 한다.**
이 실험대에는 아직 `kube-state-metrics` 가 없어 파드 readiness 가 지표로
남지 않는다. **관측 스택에 빠진 것을 이 실험이 찾아냈다** — 보완 항목이다.
---
## 8. 복구 — 자동이었다
```bash
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
```
```
재기동 시각: 11:57:09
+15초 keycloak-0 true keycloak-1 true | 외부 HTTP 200
→ 서비스 복귀
재시작 횟수: keycloak-0 = 0, keycloak-1 = 0
정지 전 세션: online 세션 5 건 살아남음
```
| | |
|---|---|
| 회복 시간 | **약 15초** (DB Ready 이후) |
| 사람 개입 | **없음** |
| Keycloak 재시작 | **불필요**`restarts=0` |
| 세션 | **살아남음** — DB 에 있으므로 |
**커넥션 풀이 스스로 재연결하고 readiness 가 다시 UP 이 되면서 Service 에
복귀했다.** `readiness` 를 쓴 설계의 이득이 여기서 나온다 — `liveness` 였다면
파드가 재시작되어 캐시까지 날아갔을 것이다.
### 개념 — readiness 와 liveness 를 가르는 기준
| | 실패하면 | 언제 쓰나 |
|---|---|---|
| **liveness** | **재시작** | 재시작하면 나아지는 문제 (교착, 메모리 누수) |
| **readiness** | **트래픽에서 격리** | 재시작해도 안 나아지는 문제 (**의존 대상이 죽음**) |
**DB 장애에 liveness 를 걸면 재앙이다.** 모든 파드가 무한 재시작하고,
DB 가 돌아와도 CrashLoopBackOff 의 백오프 때문에 회복이 늦어진다.
---
## 9. 재현 절차 (명령어)
```bash
# 0. 상주 탐침 (임시 파드는 계측에 부적합 — A-1 참조)
kubectl -n keycloak-lab run a2-probe --image=curlimages/curl:8.11.1 \
--restart=Never --command -- sleep 7200
kubectl -n keycloak-lab wait --for=condition=Ready pod/a2-probe --timeout=120s
# 1. 토큰 발급 (access 60초 안에 시험을 끝내야 한다)
kubectl -n keycloak-lab exec a2-probe -- sh -c \
'curl -s -X POST http://<k0>:8080/realms/master/protocol/openid-connect/token \
-d grant_type=password -d client_id=admin-cli \
-d username=admin -d password=<pw> > /tmp/tok.json'
# 2. DB 정지
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=90s
# 3. 네 경로
kubectl -n keycloak-lab exec a2-probe -- curl -s -o /dev/null -w '%{http_code}\n' ...
# 4. 영향 범위
kubectl -n keycloak-lab get endpoints keycloak \
-o jsonpath='{.subsets[*].addresses[*].ip}' # 비어 있으면 전면 장애
curl -s -o /dev/null -w '%{http_code}\n' https://auth.hyeonworks.com/realms/master
# 5. up 이 거짓말하는 것을 확인
curl -s "http://localhost:19090/api/v1/query?query=up%7Bjob=%22keycloak%22%7D"
# 6. 복구
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
```
---
## 10. 다음 실험에 남기는 것
| 실험 | 이 실험이 준 것 |
|---|---|
| **A-3** DB 강제 종료 | 정상 정지는 데이터를 안 잃었다. **강제 종료는?** (`synchronous_commit OFF`) |
| **A-4** 노드 상실 | postgres 가 kc-lab-2 에 있으므로 그 노드를 죽이면 **A-2 가 함께 일어난다** |
| **D-1** 백업·복구 | 여기서는 DB 가 되살아났다. **데이터가 사라졌다면?** |
| 관측 스택 | **`kube-state-metrics` 가 없어 파드 readiness 가 지표로 안 남는다** — 보완 필요 |
+323
View File
@@ -0,0 +1,323 @@
# A-3 — DB 를 강제로 죽이면 무엇을 잃는가 (RPO)
브랜치 `feature/keycloak-a3-database-crash` ·
증거 [`docs/evidence/a3-database-crash/`](evidence/a3-database-crash/) ·
2026-09-04 12:0012:05 KST · Keycloak 26.7.0 / PostgreSQL 16
선행: [`A-0`](experiment-00-session-replication.md) ·
[`A-2`](experiment-a2-database-loss.md)
---
## 0. 결론부터
```
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
그중 DB 에 실제로 존재 : 149 건
★ 유실 : 4 건
```
**로그인이 성공했다고 응답받았는데 세션이 존재하지 않는다.**
A-0 에서 발견한 `SET LOCAL synchronous_commit TO OFF` 의 대가를 실측했다.
버그가 아니라 **의도된 설계**이며, 그 비용이 얼마인지를 숫자로 확인한 것이다.
---
## 1. 설계 — 무엇을 재야 손실이 보이는가
### 1-1. `LAST_SESSION_REFRESH` 로는 못 잰다
처음 계획은 "세션 갱신 시각이 되감기는지" 보는 것이었다. 스키마를 보고 접었다.
```
created_on | integer
last_session_refresh | integer ← 초 단위
```
**손실 창은 수백 밀리초**인데 눈금이 **1초**다. 보일 리가 없다.
### 1-2. 행 존재 여부로 잰다 — 이진 판정
```
로그인 1회 = OFFLINE_USER_SESSION 행 1개
클라이언트가 sid 를 받았다 = 서버가 COMMIT 했다고 응답했다
크래시 후 그 sid 가 없다 = 잃은 것
```
**있거나 없거나**이므로 눈금 문제가 없다.
### 1-3. 그런데 로그인도 비동기 커밋인가 — **먼저 확인해야 한다**
A-0 에서 잡은 것은 **refresh** 트랜잭션이었다. 로그인(INSERT)도 그런지는
확인하지 않았다. 아니라면 이 측정 설계 자체가 성립하지 않는다.
```bash
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "alter system set log_statement='all'"
kubectl -n keycloak-lab exec deploy/postgres -- \
psql -U keycloak -d keycloak -c "select pg_reload_conf()"
```
```
BEGIN
insert into OFFLINE_USER_SESSION (...) values (...)
insert into OFFLINE_CLIENT_SESSION (...) values (...)
SET LOCAL synchronous_commit TO OFF ← 로그인도 비동기 커밋이다
COMMIT
```
**확인됐고, 함의가 refresh 보다 훨씬 무겁다.**
| | 잃으면 |
|---|---|
| refresh 갱신 시각 | 세션 수명이 조금 짧아진다. 사용자는 모른다 |
| **로그인 자체** | **토큰은 손에 있는데 세션이 없다.** 다음 요청부터 실패 |
---
## 2. 실패한 주입 ① — `--grace-period=0 --force` 는 크래시가 아니다
```bash
kubectl -n keycloak-lab delete pod -l app=postgres --grace-period=0 --force
```
```
클라이언트 성공: 291 건
DB 에 존재: 291 건
★ 유실: 0 건
```
**0건.** 그런데 이건 "안 잃었다"가 아니라 **죽인 적이 없는 것**이다.
```
=== 재기동 로그 ===
database system is ready to accept connections
(그뿐. "not properly shut down" 이 없다)
```
**crash recovery 가 돌지 않았다 = 깨끗하게 내려갔다.**
| 신호 | PostgreSQL 의 반응 |
|---|---|
| **SIGTERM** | **fast shutdown** — 진행 중 트랜잭션을 롤백하고 **WAL 을 플러시**한 뒤 종료 |
| SIGINT | smart shutdown — 연결이 끊기길 기다린다 |
| **SIGKILL** | **즉사** — 플러시 없음. 다음 기동에 crash recovery |
`--force --grace-period=0` 는 API 오브젝트를 즉시 지우지만 컨테이너 런타임은
여전히 정상 종료 절차를 밟는다. **PostgreSQL 은 SIGTERM 을 받고 얌전히
플러시했다.**
> **A-1 에서 배운 것이 또 나왔다** — 주입이 실제로 걸렸는지 먼저 확인하지
> 않으면 **"아무 일도 없었다"를 결과로 착각한다.**
> 여기서는 **crash recovery 메시지가 그 확인 수단**이다.
---
## 3. 실패한 주입 ② — 컨테이너 안에서 PID 1 은 SIGKILL 을 받지 않는다
```bash
kubectl -n keycloak-lab exec deploy/postgres -- kill -9 1
```
**아무 일도 일어나지 않았다.** 파드는 재시작하지 않았고 로그 시각도 그대로였다.
### 개념 — PID 1 의 시그널 보호
리눅스 커널은 **PID 1 을 특별 취급**한다. 자기 PID 네임스페이스 안에서 온
시그널은 **핸들러가 등록된 것만** 전달된다. **SIGKILL 도 예외가 아니다.**
```
같은 네임스페이스 안에서 → PID 1 은 등록하지 않은 시그널을 무시한다
조상 네임스페이스에서 → 전달된다 (노드에서 kill -9 하면 죽는다)
```
부팅 초기에 init 을 실수로 죽여 시스템이 멈추는 것을 막기 위한 장치인데,
컨테이너에서는 **"안에서는 PID 1 을 못 죽인다"** 로 나타난다.
---
## 4. 성공한 주입 — 백엔드 프로세스를 죽인다
PostgreSQL 은 **postmaster(부모) + 연결마다 백엔드(자식)** 구조다.
자식 하나가 비정상 종료하면 **postmaster 는 공유 메모리가 오염됐다고 보고
전체를 재초기화**한다. 그게 곧 crash recovery 다.
```bash
kubectl -n keycloak-lab exec deploy/postgres -- \
sh -c 'kill -9 $(pgrep -f "postgres: keycloak keycloak" | head -1)'
```
```
server process (PID 40) was terminated by signal 9: Killed
terminating any other active server processes
all server processes terminated; reinitializing
database system was not properly shut down; automatic recovery in progress
redo starts at 0/23CAB68
redo done at 0/2529E40
checkpoint complete: wrote 113 buffers ...
database system is ready to accept connections
```
**이번엔 주입이 걸렸다.** `not properly shut down` + `redo` 가 증거다.
파드는 재시작하지 않는다 (`restarts=0`) — 컨테이너의 PID 1 인 postmaster 는
살아 있고, 자식만 갈아치운 것이다. **데이터 관점에서는 전원이 나간 것과 같다.**
---
## 5. 결과
```
=== 크래시 전후 대조 ===
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
그중 DB 에 실제로 존재 : 149 건
★ 유실 : 4 건
=== 유실된 sid ===
★ CQUfg9HLH29xvhiu6pVlfWOo ← 토큰은 발급됐는데 세션이 없다
★ 5gLP4fqmpZBbjhH_d-0TPMMr
★ hkcOv1QskUFmYveMLB6Hljra
★ p5XybeQIYmAs818gO4Vl_5ea
```
**약 2.6% 유실.** 초당 19건 정도 로그인하던 중이었으므로
**대략 마지막 0.2초 분량**이다 — `wal_writer_delay` 기본값(200ms)과 맞는다.
### 사용자에게 어떻게 보이는가
```
로그인 성공 → access token + refresh token 을 받음
│ (크래시)
다음 요청 → access token 은 60초간 통한다
│ (서명만 보는 경로라면)
60초 후 refresh → "Session not active" → 다시 로그인
```
**즉시 드러나지 않는다.** access token 수명 동안은 정상으로 보이다가
갱신 시점에 끊긴다. 장애와 증상 사이에 **최대 60초의 시차**가 있다.
---
## 6. 개념
### WAL 과 `synchronous_commit`
```
COMMIT
├─ WAL 버퍼(메모리)에 기록 ← 항상 한다
├─ synchronous_commit = on : 디스크 플러시를 기다렸다가 응답
└─ synchronous_commit = off : 기다리지 않고 즉시 응답 ← Keycloak
└─ 크래시 시 이 구간이 사라진다
```
| 설정 | 응답 속도 | 잃는 것 |
|---|---|---|
| `on` (PostgreSQL 기본) | 느리다 (디스크 대기) | 없다 |
| **`off`** | 빠르다 | **최대 `wal_writer_delay` × 3 분량** |
**전역 설정은 `on` 이었다.**
```
전역 synchronous_commit: on
```
**Keycloak 이 자기 트랜잭션에만 `SET LOCAL` 로 끈다.** DBA 가 서버 설정만
보고 "우리는 동기 커밋"이라 믿으면 틀린다. **애플리케이션이 트랜잭션 단위로
뒤집을 수 있다.**
### crash recovery
```
기동 시 pg_control 을 읽는다
└─ "깨끗하게 종료됨" 표시가 없다
└─ "database system was not properly shut down"
└─ 마지막 체크포인트부터 WAL 을 재생(redo)
└─ 디스크에 안 내려간 커밋은 복구할 수 없다 ← 손실
```
`redo starts at 0/23CAB68``redo done at 0/2529E40` 사이가 재생된 구간이다.
**WAL 에 없는 것은 재생할 수도 없다.**
### 이 손실이 "허용된" 이유
Keycloak 의 판단은 이렇게 읽힌다.
| | |
|---|---|
| 세션 쓰기는 **매우 잦다** | 로그인마다, refresh 마다 |
| 잃어도 **회복 가능하다** | 사용자가 다시 로그인하면 된다 |
| 동기 커밋의 비용은 **모든 요청에 붙는다** | 크래시는 드물다 |
**드문 사고의 비용을 상시 지연으로 지불하지 않겠다는 선택**이다.
합리적이지만, **선택했다는 사실을 알고 있어야 한다.**
---
## 7. 운영에 주는 것
| 알게 된 것 | 함의 |
|---|---|
| 로그인도 비동기 커밋 | **RPO 가 0 이 아니다.** 크래시 시 마지막 수백 ms 로그인은 사라진다 |
| 전역 `on` 인데 세션만 `off` | **서버 설정으로 판단하면 안 된다.** 애플리케이션이 뒤집는다 |
| 손실이 즉시 안 보인다 | access token 수명만큼 시차. **모니터링은 갱신 실패율을 봐야 한다** |
| `--grace-period=0` 은 크래시가 아니다 | **장애 훈련이 훈련이 안 될 수 있다** |
| 컨테이너 안에서 PID 1 을 못 죽인다 | 크래시 재현은 **자식 프로세스**나 **노드에서** |
### 바꿀 수 있는가
```sql
-- 세션 트랜잭션까지 동기 커밋으로 강제하려면 (지연 대가를 치른다)
ALTER DATABASE keycloak SET synchronous_commit = on; -- SET LOCAL 이 이깁니다
```
**`SET LOCAL` 이 우선하므로 이것으로는 못 막는다.** Keycloak 설정이나
소스 수준의 문제이며, **RPO 0 이 필요하면 복제(streaming replication)로
푸는 것이 맞다** — 동기 스탠바이가 있으면 `synchronous_commit` 의 의미가
달라진다.
---
## 8. 재현 절차 (명령어)
```bash
# 0. 설계 확인 — 로그인도 비동기 커밋인지 먼저 본다
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "alter system set log_statement='all'" -c "select pg_reload_conf()"
# → 로그인 1회 후 로그에서 "SET LOCAL synchronous_commit TO OFF" 확인
# 1. 세션 테이블 비우기
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
-c "delete from offline_user_session"
# 2. 로그인 루프 (호스트에서 백그라운드 exec — 파드 안 & 는 exec 종료와 함께 죽는다)
kubectl -n keycloak-lab exec a2-probe -- sh -c '<로그인 반복, sid 를 /tmp/sids 에>' &
# 3. 진짜 크래시 — 백엔드 프로세스에 SIGKILL
kubectl -n keycloak-lab exec deploy/postgres -- \
sh -c 'kill -9 $(pgrep -f "postgres: keycloak keycloak" | head -1)'
# 4. 주입이 걸렸는지 확인 — 이게 없으면 결과를 해석하지 않는다
kubectl -n keycloak-lab logs deploy/postgres | grep -E "not properly shut down|redo"
# 5. 대조
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak -tAc \
"select count(*) from offline_user_session where user_session_id in (<sid 목록>)"
```
---
## 9. 다음 실험에 남기는 것
| 실험 | 이 실험이 준 것 |
|---|---|
| **D-1** 백업·복구 | RPO 는 **백업 주기 + 이 손실**이다. 둘을 더해야 진짜 RPO |
| **B-6** Redis 영속화 | `appendfsync everysec`**같은 모양의 트레이드오프** |
| **A-4** 노드 상실 | 노드가 죽으면 이것도 함께 일어난다 (postgres 가 kc-lab-2) |
| 전체 | **주입 성공 신호를 미리 정한다.** 여기서는 crash recovery 로그 |
File diff suppressed because it is too large Load Diff
+377
View File
@@ -0,0 +1,377 @@
# Keycloak 멀티노드 클러스터 — 구성과 형성 확인
로드맵 1번. 세션 저장소 실험 전부의 선행 인프라다.
브랜치 `feature/keycloak-multinode-cluster-jdbc-ping`.
**결과 — 두 파드가 서로 다른 노드에서 하나의 Infinispan 클러스터를 이뤘다.**
---
## 1. 무엇을 확인하려는가
Keycloak 26은 **디스커버리와 클러스터 통신을 서로 다른 경로로** 처리한다.
| 단계 | 경로 | 실패하면 |
|---|---|---|
| **디스커버리** — 서로를 찾는다 | PostgreSQL의 `JGROUPS_PING` 테이블 | 상대의 존재 자체를 모른다 |
| **클러스터 통신** — 실제로 대화한다 | **TCP 7800** (파드 간 직접) | **DB에는 등록되는데 클러스터가 안 붙는다** |
두 번째 줄이 이 실험대를 2노드로 만든 이유다. **단일 노드에서는 이 고장을
재현할 수 없다** — 같은 커널 안에서는 막을 경계가 없기 때문이다.
먼저 **정상적으로 붙는 상태**를 확보하고 실측값을 남긴다. 그래야 다음 실험에서
깨뜨렸을 때 무엇이 달라졌는지 비교할 수 있다.
---
## 2. 배포한 구성과 그 근거
매니페스트: [`deploy/lab/k8s/keycloak-cluster.yaml`](../deploy/lab/k8s/keycloak-cluster.yaml)
### 2-1. 왜 StatefulSet인가
Deployment를 쓰면 파드 이름이 `keycloak-7d9f8b-x4k2p`처럼 매번 바뀐다.
StatefulSet은 **`keycloak-0`, `keycloak-1`로 고정**된다.
```yaml
kind: StatefulSet
spec:
serviceName: keycloak-headless
replicas: 2
podManagementPolicy: Parallel
```
**이 실험에서 이름 안정성이 중요한 이유** — 클러스터 멤버십을 읽는 곳이 두
군데인데(Infinispan 로그, `JGROUPS_PING` 테이블) 이름이 계속 바뀌면 대조가
어렵다. 실제로 Infinispan은 `keycloak-0-49501`처럼 **파드 이름 + 랜덤 접미사**를
노드 식별자로 쓴다.
**`podManagementPolicy: Parallel`** — 기본값 `OrderedReady`는 0번이 Ready가 된
뒤에야 1번을 만든다. `Parallel`은 **동시에 시작**하므로 두 파드가 DB에 등록을
경쟁하게 되고, 그것이 운영에서 실제로 일어나는 상황이다.
### 2-2. 왜 `start`이고 `start-dev`가 아닌가
```yaml
args: ["start"]
```
`start-dev`**`cache=local`을 강제**한다. 클러스터가 아예 형성되지 않는다.
저장소의 `docker-compose.yml``start-dev`를 쓰는 것은 단일 인스턴스 학습용이며,
이 실험대에서는 쓸 수 없다.
`--optimized`는 붙이지 않았다. 붙이려면 사전 `build`가 필요하고, 없으면
첫 기동에 **암묵적 build가 실행되어 60~90초**가 걸린다. 그래서 아래처럼
`startupProbe`를 넉넉하게 준다.
### 2-3. 노드당 하나씩 배치
```yaml
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels: { app: keycloak }
```
**두 파드가 한 노드에 몰리면 7800 차단 실험이 무의미해진다.** 같은 커널 안의
루프백 통신이라 막을 대상이 없기 때문이다.
`ScheduleAnyway`를 고른 이유는 장애 실험 때문이다. `DoNotSchedule`이면 노드
하나를 죽였을 때 남은 파드가 **배치되지 못하고 Pending에 머문다.**
### 2-4. 헬스체크는 9000 포트다
```yaml
ports:
- { containerPort: 8080, name: http }
- { containerPort: 9000, name: management }
- { containerPort: 7800, name: jgroups }
startupProbe: { httpGet: { path: /health/started, port: management }, failureThreshold: 60 }
readinessProbe:{ httpGet: { path: /health/ready, port: management } }
livenessProbe: { httpGet: { path: /health/live, port: management } }
```
**Keycloak 25부터 health와 metrics가 8080이 아니라 관리 포트 9000으로 옮겨졌다.**
8080으로 프로브를 걸면 404가 나고 파드가 영원히 Ready가 되지 않는다.
`KC_HEALTH_ENABLED=true`를 켜야 엔드포인트가 노출된다.
`startupProbe``failureThreshold: 60` × `periodSeconds: 10` = **최대 10분**을
기다린다. 첫 기동의 암묵적 build 때문이다. 이게 없으면 liveness가 먼저 발동해
**재시작 루프**에 빠진다.
### 2-5. 환경변수 — 첫 실험에서 확정한 값
```yaml
- { name: KC_HOSTNAME, value: https://auth.hyeonworks.com }
- { name: KC_HOSTNAME_STRICT, value: "true" }
- { name: KC_PROXY_HEADERS, value: xforwarded }
- { name: KC_HTTP_ENABLED, value: "true" }
```
[`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md)에서
측정으로 확정한 조합이다.
| 설정 | 역할 |
|---|---|
| `KC_HOSTNAME`**전체 URL** | 스킴·호스트를 **고정**한다. 헤더와 무관하게 `iss`가 https로 발급된다 |
| `KC_HOSTNAME_STRICT=true` | Host 헤더를 믿지 않는다. 조작으로 흐름을 돌릴 여지를 없앤다 |
| `KC_PROXY_HEADERS=xforwarded` | **클라이언트 IP** 등 나머지를 forwarded 헤더에서 가져온다 |
| `KC_HTTP_ENABLED=true` | 앞단이 TLS를 끊었으므로 평문 HTTP를 받는다 |
**이 실험을 먼저 하지 않았다면** 지금 `iss``http://10.42.x.x`로 나왔을 것이고,
원인을 세션 쪽에서 찾느라 헤맸을 것이다.
### 2-6. 힙 상한
```yaml
- { name: JAVA_OPTS_KC_HEAP, value: "-Xms256m -Xmx512m" }
resources:
requests: { memory: 640Mi, cpu: 100m }
limits: { memory: 900Mi }
```
Keycloak은 기본값이 넉넉해 그냥 두면 1GB를 넘긴다. 이 실험대의 게스트 여유가
약 3.8GB이므로 명시적으로 잡는다. 실측 결과 **파드당 약 590Mi**로 안정됐다.
### 2-7. PostgreSQL — 볼륨이 노드에 고정된다
```yaml
storageClassName: local-path
strategy:
type: Recreate
env:
- { name: PGDATA, value: /var/lib/postgresql/data/pgdata }
```
k3s 기본 `local-path` 프로비저너는 **파드가 배치된 노드의 로컬 디스크**에
볼륨을 만든다. 따라서 PostgreSQL은 그 노드에 묶인다.
**이것은 결함이 아니라 실험 조건이다.** 나중에 "데이터베이스가 있는 노드가
죽으면" 시나리오가 그래서 의미를 갖는다.
- `strategy: Recreate` — RWO 볼륨은 두 파드가 동시에 마운트할 수 없다.
기본값 `RollingUpdate`면 새 파드가 볼륨을 못 잡고 멈춘다
- `PGDATA`를 한 단계 아래로 — 마운트 지점에 `lost+found` 같은 것이 있으면
`initdb`가 거부한다
### 2-8. 헤드리스 서비스는 왜 두는가
```yaml
kind: Service
metadata: { name: keycloak-headless }
spec:
clusterIP: None
```
**jdbc-ping 디스커버리에는 필요 없다.** DB로 서로를 찾기 때문이다.
개별 파드에 안정된 DNS 이름으로 접근해 상태를 조회하기 위해 둔다.
---
## 3. 실행한 명령
### 3-1. 브랜치와 정리
```bash
# 워크스테이션
cd ~/workspace/keycloak-pattern
git checkout -b feature/keycloak-multinode-cluster-jdbc-ping
git merge --no-edit develop-keycloak-session-store
# lab host — 끝난 실험을 지워 메모리를 회수한다
kubectl delete ns header-lab
```
정리 후 게스트 사용량이 `kc-lab-1 1593Mi(46%)` / `kc-lab-2 872Mi(35%)`로 떨어졌다.
### 3-2. 배포
```bash
# 워크스테이션 — 매니페스트 작성 후
git add deploy/lab/k8s/keycloak-cluster.yaml
git commit -m "feat: deploy Keycloak multi-node cluster with PostgreSQL"
git push -u origin feature/keycloak-multinode-cluster-jdbc-ping
# lab host
cd ~/workspace/keycloak-pattern
git fetch origin
git checkout -b feature/keycloak-multinode-cluster-jdbc-ping origin/feature/keycloak-multinode-cluster-jdbc-ping
kubectl apply -f deploy/lab/k8s/keycloak-cluster.yaml
```
**PostgreSQL을 먼저 기다린다.** Keycloak이 DB 없이 뜨면 기동에 실패한다.
```bash
kubectl -n keycloak-lab rollout status deployment/postgres --timeout=180s
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=600s
```
이미지를 당겨오고 암묵적 build가 도는 첫 기동은 **수 분** 걸린다.
### 3-3. 검증
```bash
# 파드 배치 — 서로 다른 노드에 있어야 한다
kubectl -n keycloak-lab get pods -o wide
# 클러스터 뷰 — Infinispan 로그
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
# 디스커버리 테이블
PG=$(kubectl -n keycloak-lab get pod -l app=postgres -o name | head -1)
kubectl -n keycloak-lab exec "$PG" -- \
psql -U keycloak -d keycloak -c "SELECT name, cluster_name, ip, coord FROM jgroups_ping ORDER BY name;"
# 외부 접근과 issuer
curl -s https://auth.hyeonworks.com/realms/master/.well-known/openid-configuration | python3 -m json.tool
# 자원
kubectl -n keycloak-lab top pods
```
---
## 4. 확인된 사실
증거 원자료: [`evidence/keycloak-multinode-cluster/`](evidence/keycloak-multinode-cluster/)
### 4-1. 클러스터가 형성됐다
```
ISPN000094: Received new cluster view for channel ISPN:
[keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938, keycloak-0-49501]
↑ 멤버 수
ISPN100000: Node keycloak-0-49501 joined the cluster
ISPN000079: Channel `ISPN` local address is `keycloak-0-49501`,
physical addresses are `[10.42.1.18:7800]`
```
두 파드가 **동일한 뷰**를 보고 있고, 물리 주소가 **7800**임이 로그에 찍힌다.
### 4-2. 디스커버리와 통신이 분리되어 있다
```
name | cluster_name | ip | coord
------------------+--------------+-----------------+-------
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
```
**테이블 하나에 두 메커니즘이 다 보인다.**
- `name`·`cluster_name` — **DB로 하는 디스커버리**의 결과
- `ip` 컬럼의 `:7800`**실제 통신이 일어날 경로**
`coord``t``keycloak-1`이 코디네이터다. 이 노드를 죽였을 때 인계가
일어나는지가 다음 실험 항목이다.
전체 스키마는 `address / name / cluster_name / ip / coord / last_update /
coordinated_by`이며 기본키는 `address`다.
### 4-3. 노드당 하나씩 배치됐다
```
keycloak-0 10.42.1.18 kc-lab-2
keycloak-1 10.42.0.16 kc-lab-1
postgres 10.42.1.19 kc-lab-2
```
파드 IP 대역이 노드를 알려준다(`10.42.0.x` = kc-lab-1, `10.42.1.x` = kc-lab-2).
**독립된 커널 두 개에 하나씩** 떴으므로 7800 차단 실험의 전제가 성립한다.
PostgreSQL이 `kc-lab-2`에 있다는 점도 기록해둔다. **`kc-lab-2`를 죽이면
Keycloak 하나와 데이터베이스가 동시에 사라진다.**
### 4-4. 2홉 헤더 계약이 실제로 작동한다
```
issuer https://auth.hyeonworks.com/realms/master
authorization_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/auth
token_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/token
end_session_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/logout
jwks_uri https://auth.hyeonworks.com/realms/master/protocol/openid-connect/certs
```
**전부 `https`이고 외부 호스트명이다.** 첫 실험의 결론이 그대로 값을 했다.
### 4-5. 자원
```
keycloak-0 594Mi
keycloak-1 593Mi
postgres 67Mi
──────────────────────
kc-lab-1 2248Mi (65%)
kc-lab-2 1447Mi (58%)
```
예상(파드당 700Mi)보다 적다. `JAVA_OPTS_KC_HEAP` 제한이 작동했다.
BFF와 Redis를 추가할 여유가 남아 있다.
---
## 5. 겪은 함정
### `JGROUPS_PING` 컬럼명은 자료마다 다르다
오래된 문서에는 `own_addr`, `ping_data` 같은 이름이 나오지만 **Keycloak 26의
실제 스키마는 다르다.**
```
address / name / cluster_name / ip / coord / last_update / coordinated_by
```
쿼리 전에 `\d jgroups_ping`으로 확인한다.
### Keycloak 컨테이너에 `curl`이 없다
메트릭을 파드 안에서 조회하려다 실패했다.
```
sh: line 1: curl: command not found
```
Keycloak 공식 이미지는 최소 구성이다. 메트릭을 볼 때는 포트포워딩하거나
임시 파드를 쓴다.
```bash
kubectl -n keycloak-lab port-forward keycloak-0 9000:9000 &
curl -s localhost:9000/metrics | grep -i cluster
# 또는
kubectl -n keycloak-lab run m --rm -i --restart=Never --image=curlimages/curl:8.11.1 -- \
curl -s http://keycloak-0.keycloak-headless:9000/metrics
```
### 첫 기동이 느린 것은 정상이다
`--optimized` 없이 `start`하면 **암묵적 build**가 실행된다. `startupProbe`
넉넉히 주지 않으면 liveness가 먼저 발동해 재시작 루프에 빠진다.
---
## 6. 다음 실험 — 깨뜨려서 무엇이 보이는지
정상 상태를 확보했으므로 이제 의도적으로 고장을 만든다.
| 실험 | 방법 | 확인할 것 |
|---|---|---|
| **7800 차단** | NetworkPolicy로 파드 간 7800만 차단 | **DB엔 등록되는데 클러스터가 안 붙는** 증상. 로그에 무엇이 먼저 보이는가 |
| **노드 상실** | `virsh destroy kc-lab-2` | 코디네이터 인계가 일어나는가. PostgreSQL도 같이 죽는다는 점에 유의 |
| **DB 상실** | postgres 파드 정지 | 이미 형성된 클러스터는 버티는가. 새 로그인은? |
**7800 차단부터 하는 것이 좋다.** 되돌리기가 가장 쉽고(NetworkPolicy 삭제),
증상이 로그에 선명하게 남는다.
## 참고
| 문서 | 관계 |
|---|---|
| [`session-store-lab-roadmap.md`](session-store-lab-roadmap.md) | 이 실험은 로드맵 1번 |
| [`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md) | `KC_HOSTNAME`·`KC_PROXY_HEADERS` 값의 근거 |
| [`session-lab-operations.md`](session-lab-operations.md) | 명령·자원 예산 |
| [`session-lab-concepts.md`](session-lab-concepts.md) | StatefulSet·프로브·PVC 등 개념 |
+429
View File
@@ -0,0 +1,429 @@
# 관측성 — Prometheus · node-exporter · Grafana
로드맵 10번. 장애 주입 실험보다 **먼저** 세운다.
**왜 먼저인가** — 나중에 세우면 이미 지나간 장애의 지표를 볼 수 없다.
"클러스터가 1분쯤 뒤에 복구됐다"는 측정이 아니라 인상이다.
로드맵에 *"장애 주입 중에 어떤 지표가 먼저 움직이는지 기록한다"*고 적어둔 항목은
관측이 먼저 서 있어야만 가능하다.
메모리를 8GB → 12GB로 증설한 뒤에야 올릴 수 있게 됐다.
---
## 1. 무엇을 세웠나
```
┌─ Grafana ──────────┐
브라우저 ──────▶│ app2.hyeonworks.com│ 대시보드
└─────────┬──────────┘
│ PromQL
┌─────────▼──────────┐
│ Prometheus │ 수집·저장 (TSDB, 7일)
└─────────┬──────────┘
│ scrape (15초)
┌───────────────────┼───────────────────┐
▼ ▼ ▼
Keycloak :9000 node-exporter :9100 kubelet
(앱 지표) (머신 지표) (컨테이너 지표)
```
매니페스트: [`deploy/lab/k8s/observability.yaml`](../deploy/lab/k8s/observability.yaml)
| 구성요소 | 역할 | 실측 메모리 |
|---|---|---|
| Prometheus | 수집·저장·질의 | 164Mi |
| node-exporter (DaemonSet) | 노드당 하나, 머신 지표 | 8Mi × 2 |
| Grafana | 시각화 | 65Mi |
| **합계** | | **약 245Mi** |
예상(550Mi)보다 훨씬 적다. 실험대 규모에서는 관측성 비용이 거의 무시할 수준이다.
---
## 2. 왜 kube-prometheus-stack을 쓰지 않았나
Helm 차트 하나로 끝내는 방법이 있지만 **평범한 매니페스트를 직접 썼다.**
| | kube-prometheus-stack | 직접 작성 |
|---|---|---|
| 설치 | Helm 한 줄 | 매니페스트 400줄 |
| 메모리 | 1.5GB 이상 | **245Mi** |
| 포함 | Operator, Alertmanager, 대시보드 다수, kube-state-metrics | 필요한 것만 |
| **보이는 것** | 추상화 뒤에 숨음 | **스크레이프 설정·RBAC·relabel 이 눈에 보임** |
세 번째 줄이 결정적이다. 이 실험대의 목적은 **인과를 직접 확인하는 것**이므로,
"어떻게 타깃을 찾는가"가 YAML에 드러나 있어야 한다. Operator를 쓰면
`ServiceMonitor` 하나만 보이고 그 아래는 감춰진다.
---
## 3. 구성 결정과 근거
### 3-1. 관측 스택의 배치 — 장애 도메인 분리
```yaml
nodeSelector:
node-role.kubernetes.io/control-plane: "true"
```
**관측 시스템은 관측 대상과 같은 장애 도메인에 있으면 안 된다.** 죽는 순간을
기록해야 하는데 같이 죽으면 기록이 남지 않는다.
노드가 둘뿐이라 완전히 피할 수는 없다. 그래서 규칙을 정했다.
| 노드 | 역할 | 실험에서 |
|---|---|---|
| **kc-lab-1** (k3s **server**) | control plane · Traefik · coredns · metrics-server · local-path-provisioner | **관측 스택을 여기 둔다. 죽이지 않는다** |
| **kc-lab-2** (k3s **agent**) | keycloak-0 · postgres | **장애 주입 대상** |
`kubernetes.io/hostname`으로 못박지 않고 **`node-role.kubernetes.io/control-plane`
라벨**을 쓴 이유는 의미가 드러나기 때문이다 — "컨트롤 플레인 노드에 둔다"는
의도가 호스트 이름보다 오래간다.
### 3-2. 앞선 판단을 정정했다
배치를 조사하기 전에는 **"노드 상실 실험은 `kc-lab-1`을 죽여서 하자"**고
적었다. 그 노드에 Keycloak 하나만 있다고 생각했기 때문이다. **틀렸다.**
```
kc-lab-1 (server) keycloak-1, traefik, coredns, metrics-server, local-path-provisioner
kc-lab-2 (agent) keycloak-0, postgres
```
`kc-lab-1`을 죽이면 **API 서버·DNS·인그레스가 한꺼번에 사라진다.** 노드 상실이
아니라 **컨트롤 플레인 상실**이며, `kubectl`조차 동작하지 않는다.
**깨끗한 워커 노드 상실 실험은 `kc-lab-2`를 죽이는 것이다.** 그때도 변수가
둘(keycloak-0 + postgres)이지만, 클러스터 제어는 살아 있고 관측도 계속된다.
### 3-3. 스크레이프 주기 15초
```yaml
global:
scrape_interval: 15s
```
운영에서는 30~60초가 흔하지만 여기서는 짧게 잡았다. **노드가 죽는 순간을
두어 샘플 안에 잡아야** "무엇이 먼저 움직였나"를 말할 수 있다.
60초면 장애와 복구가 같은 샘플에 뭉개진다.
### 3-4. 타깃을 정적 목록으로 두지 않는다
```yaml
kubernetes_sd_configs:
- role: endpoints
namespaces: { names: [keycloak-lab] }
```
**파드 IP는 재시작마다 바뀐다.** 실험대를 전원 종료했다 켰을 때 모든 파드가
새 주소를 받는 것을 직접 확인했다(`10.42.1.22``10.42.1.25`).
정적 목록을 적어두면 그때마다 깨진다.
쿠버네티스 API에 물어보는 방식(service discovery)이므로 **파드가 옮겨다녀도
따라간다.** Traefik의 `trustedIPs`에 개별 IP를 적을 수 없었던 것과 같은 이유다.
### 3-5. relabel — 발견한 것을 걸러내고 이름을 붙인다
```yaml
relabel_configs:
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
action: keep
regex: keycloak-headless;management
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
- source_labels: [__meta_kubernetes_pod_node_name]
target_label: node
```
service discovery는 네임스페이스의 **모든 엔드포인트**를 가져온다. 그중
필요한 것만 남기고 나머지는 버리는 것이 `keep`이다.
- 첫 규칙 — `keycloak-headless` 서비스의 `management` 포트만 남긴다.
8080(http)까지 긁으면 애플리케이션 트래픽 포트에 헛되이 요청이 간다
- 나머지 두 규칙 — **`pod``node` 라벨을 붙인다.** 이것이 없으면
"어느 파드가, 어느 노드에서" 라는 질문에 답할 수 없다.
노드 상실 실험에서 결정적이다
### 3-6. Keycloak 지표는 9000 포트다
헬스체크와 같은 관리 포트다. `KC_METRICS_ENABLED=true`가 이미 StatefulSet에
설정돼 있다. **8080을 긁으면 지표가 나오지 않는다.**
### 3-7. node-exporter는 DaemonSet + 호스트 네임스페이스
```yaml
kind: DaemonSet
spec:
template:
spec:
hostNetwork: true
hostPID: true
tolerations:
- operator: Exists
```
- **DaemonSet** — 노드마다 정확히 하나. 죽을 노드에도 있어야 **꺼지기 직전의
마지막 샘플**이 남는다
- **`hostNetwork`/`hostPID`** — 측정 대상이 컨테이너가 아니라 **머신**이다.
컨테이너 네임스페이스 안에서 보면 자기 자신만 보인다
- **`tolerations: operator: Exists`** — 어떤 taint가 걸린 노드에도 뜬다.
관측이 빠지는 노드가 있으면 안 된다
### 3-8. Prometheus 저장소는 PVC
```yaml
storageClassName: local-path
--storage.tsdb.retention.time=7d
```
`emptyDir`로 두면 파드가 재시작될 때 **장애 실험의 기록이 통째로 사라진다.**
사후 추적이 목적이므로 영속 저장이 필요하다.
`local-path`는 노드에 고정되므로 Prometheus도 `kc-lab-1`에 묶인다.
`nodeSelector`와 방향이 같아 문제가 되지 않는다.
보존 7일은 실험 기간보다 넉넉하면서 **볼륨이 노드를 채우는 원인이 되지 않을**
크기다.
```yaml
securityContext:
fsGroup: 65534
```
`prom/prometheus` 이미지는 `nobody`(65534)로 실행된다. `fsGroup`이 없으면
새로 만들어진 볼륨의 소유자가 root라 **쓰기 권한이 없어 기동에 실패한다.**
### 3-9. Grafana에도 외부 URL을 알려줘야 한다
```yaml
- name: GF_SERVER_ROOT_URL
value: https://app2.hyeonworks.com
```
**Keycloak의 `KC_HOSTNAME`과 정확히 같은 성격의 설정이다.** Grafana도
리다이렉트와 자산 경로에 절대 URL을 만든다. 이 값이 없으면 로그인 리다이렉트가
`http://<파드IP>:3000`으로 나간다.
2홉 헤더 계약에서 확인한 원리가 여기서도 그대로 적용된다 —
**프록시 뒤의 애플리케이션은 자기가 외부에서 어떤 주소로 보이는지 모른다.**
### 3-10. 데이터소스는 파일로 프로비저닝
```yaml
volumeMounts:
- name: datasources
mountPath: /etc/grafana/provisioning/datasources
```
UI에서 클릭으로 추가하면 Grafana 자체 DB에만 남는다. 그 DB는 여기서
`emptyDir`이므로 **파드가 재시작되면 사라진다.** 파일로 두면 항상 같은 상태로
뜬다.
### 3-11. Grafana를 `app2`에 붙인 이유
인증서에 들어 있는 이름이 `auth` / `app1` / `app2` 셋뿐이고 `app2`가 비어
있었다. **SSO 실험에서 `app2`가 필요해지면 옮긴다.**
---
## 4. 실행한 명령
```bash
# 워크스테이션 — 매니페스트 작성 후
git add deploy/lab/k8s/observability.yaml
git commit -m "feat: add Prometheus, node-exporter and Grafana"
git push origin feature/keycloak-multinode-cluster-jdbc-ping
# lab host
cd ~/workspace/keycloak-pattern && git pull
kubectl apply -f deploy/lab/k8s/observability.yaml
kubectl -n observability rollout status deployment/prometheus --timeout=300s
kubectl -n observability rollout status daemonset/node-exporter --timeout=180s
kubectl -n observability rollout status deployment/grafana --timeout=300s
```
**검증 — 배포 성공과 타깃 수집은 다른 문제다.**
```bash
kubectl -n observability run q --rm -i --restart=Never \
--image=curlimages/curl:8.11.1 --quiet --command -- \
curl -s "http://prometheus.observability.svc:9090/api/v1/targets?state=any" > /tmp/targets.json
python3 -c "
import json
d,_ = json.JSONDecoder().raw_decode(open('/tmp/targets.json').read())
ts = d['data']['activeTargets']
print(f\"{sum(1 for t in ts if t['health']=='up')}/{len(ts)} up\")
for t in ts:
if t['health'] != 'up': print(t['labels'], t.get('lastError'))
"
```
---
## 5. 겪은 함정
### kubelet 타깃이 403 Forbidden
첫 배포에서 **7개 중 5개만 up**이었다.
```
DOWN kubelet kc-lab-1 server returned HTTP status 403 Forbidden
DOWN kubelet kc-lab-2 server returned HTTP status 403 Forbidden
```
원인은 RBAC였다. kubelet 지표는 **API 서버의 proxy 서브리소스**를 통해
가져온다.
```
/api/v1/nodes/<name>/proxy/metrics
─────
```
이 경로에는 `nodes``nodes/metrics`가 아니라 **`nodes/proxy`** 권한이
필요하다.
```diff
- resources: [nodes, nodes/metrics, services, endpoints, pods]
+ resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
```
**다른 잡은 전부 정상이었다.** 이런 부분 실패는 타깃 목록을 직접 확인하지
않으면 드러나지 않는다. `rollout status`는 "성공"이라고 말한다.
### `kubectl run --rm -i`의 출력에 종료 메시지가 섞인다
```
json.decoder.JSONDecodeError: Extra data: line 1 column 54973
```
`kubectl run --rm`은 컨테이너 출력 뒤에 `pod "q" deleted`를 덧붙인다.
JSON 파서가 그 뒤를 만나면 실패한다.
**해결**`raw_decode`로 앞쪽의 완전한 JSON만 읽는다.
```python
d, _ = json.JSONDecoder().raw_decode(raw)
```
---
## 6. 실험에 쓸 지표
메트릭 이름이 **1506개** 수집된다. 그중 장애 실험에서 볼 것들이다.
### 가장 중요한 것 — `up`
```promql
up
up{job="keycloak"}
```
Prometheus가 타깃을 긁는 데 성공했는가를 0/1로 알려주는 **합성 지표**다.
타깃이 응답하지 않으면 0이 된다.
**노드나 파드가 죽는 순간 가장 먼저 움직이는 신호**이며, 다른 모든 지표가
사라지는 것과 달리 `up`**0이라는 값으로 남는다.** 그래서 "언제부터 죽었나"를
사후에 알 수 있다.
### JGroups — 7800 차단 실험의 핵심
```promql
vendor_jgroups_fd_sock2_get_num_suspected_members
vendor_jgroups_merge3_get_views
vendor_jgroups_tcp_get_different_cluster_messages
```
| 지표 | 무엇을 말하는가 |
|---|---|
| `fd_sock2_..._suspected_members` | **FD_SOCK2가 의심하는 멤버 수.** 현재 두 파드 모두 `0`. 7800이 막히면 상대를 suspect 하기 시작한다 |
| `merge3_get_views` | **MERGE3가 처리한 뷰 수.** split brain 후 다시 합칠 때 움직인다 |
| `tcp_get_different_cluster_messages` | 다른 클러스터로부터 온 메시지 |
**7800 차단 실험의 가설**`JGROUPS_PING` 테이블은 그대로 채워진 채
`suspected_members`가 0에서 1로 오르고, 클러스터 뷰가 각각 1로 쪼개진다.
### 노드 지표
```promql
node_memory_MemAvailable_bytes
node_load1
node_network_receive_bytes_total
node_filesystem_avail_bytes
```
**"머신이 죽었나 프로세스가 죽었나"** 를 가르는 데 쓴다. 파드는 사라졌는데
node-exporter가 살아 있으면 프로세스 문제이고, 둘 다 사라지면 머신 문제다.
### Keycloak 애플리케이션 지표
```promql
keycloak_session_expiration_task_seconds_count
```
`keycloak_` 접두 지표는 아직 적다. 세션 관련 지표는 **실제 로그인이 발생해야**
나타나므로, 세션 복제 실험 이후 다시 조사한다.
---
## 7. 접근
| | 주소 | 계정 |
|---|---|---|
| Grafana | `https://app2.hyeonworks.com` | `admin` / `lab-grafana-change-me` |
| Prometheus | 클러스터 내부 `prometheus.observability.svc:9090` | — |
Prometheus UI를 직접 보려면 포트포워딩한다.
```bash
kubectl -n observability port-forward svc/prometheus 9090:9090
# http://localhost:9090/targets
```
**Grafana 비밀번호가 매니페스트에 평문이다.** 로드맵 11번(비밀 관리)에서
정리한다. 지금 드러내 두는 것은 의도이며, 감춰두면 잊어버린다.
---
## 8. 자원 실측
```
grafana 65Mi
prometheus 164Mi
node-exporter 8Mi × 2
────────────────────────
합계 약 245Mi
kc-lab-1 2045Mi (41%)
kc-lab-2 1131Mi (28%)
호스트 여유 3957MB
```
메모리 증설(8GB → 12GB) 전이었다면 kc-lab-1이 60%를 넘겼을 것이다.
증설이 이 항목을 가능하게 했다.
---
## 9. 다음
관측이 서 있으므로 이제 고장을 주입하면 **무엇이 먼저 움직였는지**가 기록된다.
```
0. 세션 복제 확인 ← 로그인 세션을 만들어 두 노드에 복제되는지
1. TCP 7800 차단 ← suspected_members 와 JGROUPS_PING 대조
2. DB 상실 ← postgres 파드 정지
3. 노드 상실 ← kc-lab-2 (agent) 를 죽인다. kc-lab-1 이 아니다
```
각 실험 전후로 같은 PromQL을 실행해 대조한다.
## 참고
| 문서 | 관계 |
|---|---|
| [`keycloak-multinode-cluster.md`](keycloak-multinode-cluster.md) | 관측 대상의 구성 |
| [`session-lab-concepts.md`](session-lab-concepts.md) | Prometheus·RBAC·DaemonSet 등 개념 |
| [`session-store-lab-roadmap.md`](session-store-lab-roadmap.md) | 로드맵 10번 |
| [`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md) | `GF_SERVER_ROOT_URL`이 필요한 이유 |
+10 -2
View File
@@ -1,7 +1,15 @@
# 열린 질문 커버리지 — 이 실험대로 답할 수 있는가 # 열린 질문 커버리지 — 이 실험대로 답할 수 있는가
공개 기록(`hyeonworks.com/questions`)에 등록된 KeyCloak Patterns 열린 질문 공개 기록에 등록된 KeyCloak Patterns 열린 질문 네 개를, 이 실험대가 실제로
네 개를, 이 실험대가 실제로 검증할 수 있는지 대조한 결과. 검증할 수 있는지 대조한 결과.
> **목록 경로는 [`/explore/questions`](https://hyeonworks.com/explore/questions)**
> 다. `/questions` 는 404 이고 개별 문서만 `/questions/<slug>` 로 열린다.
>
> **2026-09-04 재확인** — Playwright 로 네 문서를 전문 재독하고
> 「남은 미지수」·「다음 검증」·「제약」을 항목 단위로 대조한 결과
> **계획에 빠진 항목 9개**를 찾아 보강했다. 항목별 실험 번호 대조표는
> [`experiment-plan.md`](experiment-plan.md) B층 머리에 있다.
**결론 — 네 개 모두 이 실험대에서 재현 가능하다. 다만 로드맵에 빠진 항목이 **결론 — 네 개 모두 이 실험대에서 재현 가능하다. 다만 로드맵에 빠진 항목이
있고, 순서가 한 곳 뒤집혀 있다.** 있고, 순서가 한 곳 뒤집혀 있다.**
+597 -8
View File
@@ -2826,17 +2826,606 @@ SSH 공개키 두 줄이다. 공개키 자체는 비밀이 아니지만, **저
--- ---
## 10층. 쿠버네티스 리소스 — 이 실험대에서 실제로 쓴 것들
5층이 k3s 자체라면 여기는 그 위에 올린 리소스들이다.
### 워크로드 세 종류 — 무엇을 언제 쓰는가
| | 보장하는 것 | 이 실험대에서 |
|---|---|---|
| **Deployment** | 파드 N개를 유지. 이름은 매번 바뀐다 | postgres, grafana, prometheus, echo |
| **StatefulSet** | **안정된 이름**(`-0`, `-1`)과 순서 | **keycloak** |
| **DaemonSet** | **노드마다 정확히 하나** | node-exporter, svclb |
**StatefulSet을 Keycloak에 쓴 이유** — Infinispan이 **파드 이름 + 랜덤 접미사**를
클러스터 노드 식별자로 쓴다(`keycloak-0-49501`). Deployment면 이름이
`keycloak-7d9f8b-x4k2p`처럼 매번 달라져서, 로그와 `JGROUPS_PING` 테이블을
대조하기가 어려워진다.
**`podManagementPolicy`**
| 값 | 동작 |
|---|---|
| `OrderedReady` (기본) | `-0`이 Ready가 된 뒤에야 `-1`을 만든다 |
| **`Parallel`** | **동시에 시작한다** |
이 실험대는 `Parallel`을 쓴다. 두 파드가 **동시에 클러스터 등록을 시도하는 것**이
운영에서 실제로 일어나는 상황이기 때문이다.
**DaemonSet을 node-exporter에 쓴 이유** — replica 수를 지정하지 않는다.
노드가 늘면 자동으로 늘고, 줄면 준다. **죽을 노드에도 반드시 있어야**
꺼지기 직전의 마지막 샘플이 남는다.
```bash
kubectl get deploy,sts,ds -A
```
### 저장소 — PVC · PV · StorageClass
```
PersistentVolumeClaim (PVC) "5Gi 짜리 읽기쓰기 볼륨을 주세요" ← 요청
│ storageClassName: local-path
StorageClass 어떻게 만들지 아는 프로비저너
PersistentVolume (PV) 실제로 만들어진 볼륨 ← 결과
```
**PVC는 요청서, PV는 실물이다.** 파드는 PVC 이름만 알면 되고, 그 뒤가
로컬 디스크인지 NFS인지 클라우드 블록 스토리지인지 몰라도 된다.
**`accessModes`**
| 값 | 의미 |
|---|---|
| **`ReadWriteOnce` (RWO)** | **한 노드에서만** 읽기/쓰기 |
| `ReadOnlyMany` | 여러 노드에서 읽기만 |
| `ReadWriteMany` | 여러 노드에서 읽기/쓰기 (NFS 등) |
**RWO가 `strategy: Recreate`를 강제한다.** 기본값 `RollingUpdate`는 새 파드를
띄운 뒤 옛 파드를 내리는데, RWO 볼륨은 **두 파드가 동시에 마운트할 수 없어서**
새 파드가 영원히 Pending에 머문다.
```yaml
strategy:
type: Recreate # 옛 파드를 먼저 내리고 새 파드를 띄운다
```
**k3s의 `local-path` 프로비저너 — 볼륨이 노드에 못박힌다**
```json
"nodeAffinity": {
"required": { "nodeSelectorTerms": [{
"matchExpressions": [{ "key": "kubernetes.io/hostname", "values": ["kc-lab-2"] }]
}]}
}
경로: /var/lib/rancher/k3s/storage/pvc-<uuid>_<ns>_<name>
```
**그 노드의 로컬 디스크에 디렉터리를 만드는 것이 전부**다. 따라서
**PVC를 쓰는 파드는 그 노드를 벗어날 수 없다.**
| 결과 | |
|---|---|
| 노드가 죽으면 | **파드가 다른 노드로 재배치되지 못한다** |
| 실험 관점 | **결함이 아니라 조건이다.** "DB가 있는 노드가 죽으면"이 의미를 갖는다 |
```bash
kubectl get pvc -A
kubectl get pv
kubectl get pv <name> -o jsonpath='{.spec.nodeAffinity}' | python3 -m json.tool
```
### Secret — 감춰지지 않는다
```yaml
kind: Secret
type: Opaque
stringData:
POSTGRES_PASSWORD: lab-postgres-change-me
```
`stringData`는 평문으로 쓰고 쿠버네티스가 base64로 인코딩해 저장한다.
`data`는 직접 base64로 넣는다.
**base64는 암호화가 아니라 인코딩이다.**
```bash
kubectl -n keycloak-lab get secret keycloak-lab-secrets -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
```
한 줄로 읽힌다. etcd에도 그대로 들어 있다.
| 그래도 Secret을 쓰는 이유 | |
|---|---|
| RBAC로 접근을 나눌 수 있다 | ConfigMap과 별도로 권한 관리 |
| 로그·`describe`에 값이 안 찍힌다 | 사고로 노출될 확률이 준다 |
| 볼륨·env 주입 방식이 표준화된다 | |
**진짜 보호는 별도 계층이다** — SealedSecret, 외부 KMS, 또는 클라우드
시크릿 매니저. 로드맵 11번의 주제다.
### RBAC — ServiceAccount · ClusterRole · Binding
Prometheus가 쿠버네티스 API에 물어서 타깃을 찾으려면 **읽기 권한**이 필요하다.
```
ServiceAccount 파드가 쓰는 신원 (누구인가)
ClusterRoleBinding 신원과 권한을 잇는다
ClusterRole 무엇을 할 수 있는가 (리소스 × 동사)
```
```yaml
rules:
- apiGroups: [""]
resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
verbs: [get, list, watch]
```
**`Role`과 `ClusterRole`의 차이** — `Role`은 한 네임스페이스 안에서만,
`ClusterRole`은 클러스터 전체에서 유효하다. 노드는 네임스페이스에 속하지
않으므로 **노드를 읽으려면 반드시 `ClusterRole`**이다.
**서브리소스가 따로 있다 — 실제로 걸린 함정**
`nodes`, `nodes/metrics`, `nodes/proxy`는 **서로 다른 권한**이다.
```
/api/v1/nodes/<name>/proxy/metrics
─────
이 경로에는 nodes/proxy 가 필요
```
`nodes/proxy`를 빠뜨렸을 때 kubelet 타깃만 **403 Forbidden**으로 실패하고
나머지 잡은 전부 정상이었다. **부분 실패라 `rollout status`는 성공이라고
말한다.** 타깃 목록을 직접 봐야 드러난다.
```bash
kubectl auth can-i get nodes/proxy --as=system:serviceaccount:observability:prometheus
kubectl describe clusterrole prometheus
```
### 배치 제어 — nodeSelector · 라벨 · taint
```yaml
nodeSelector:
node-role.kubernetes.io/control-plane: "true"
```
**호스트 이름 대신 역할 라벨을 쓴다.** `kubernetes.io/hostname: kc-lab-1`로
못박으면 노드 이름이 바뀔 때 깨지고, **왜 거기 두는지가 드러나지 않는다.**
k3s는 server 노드에 `node-role.kubernetes.io/control-plane=true`를 붙인다.
```bash
kubectl get nodes --show-labels
kubectl get nodes -l node-role.kubernetes.io/control-plane=true
```
**taint와 toleration**
| | |
|---|---|
| **taint** | 노드에 붙는 "여기 오지 마" 표시 |
| **toleration** | 파드가 갖는 "그래도 갈 수 있음" 면제권 |
```yaml
tolerations:
- operator: Exists # 어떤 taint 든 무시한다
```
node-exporter에 이걸 주는 이유는 **관측이 빠지는 노드가 있으면 안 되기**
때문이다. taint가 걸린 노드에서도 떠야 한다.
**배치를 정하는 세 수단의 차이**
| 수단 | 성격 |
|---|---|
| `nodeSelector` | **반드시** 그 라벨의 노드에 |
| `topologySpreadConstraints` | **골고루** 퍼뜨린다 |
| taint / toleration | 노드가 **거부**하고 파드가 **면제**받는다 |
### k3s server와 agent — 죽였을 때가 다르다
```bash
kubectl get nodes -o custom-columns=\
'NODE:.metadata.name,CP:.metadata.labels.node-role\.kubernetes\.io/control-plane'
```
| | kc-lab-1 (**server**) | kc-lab-2 (**agent**) |
|---|---|---|
| 실행 | API 서버 · 스케줄러 · etcd(SQLite) | kubelet · containerd |
| 이 실험대에서 | keycloak-1 · traefik · **coredns** · metrics-server · local-path-provisioner | keycloak-0 · postgres |
| 죽이면 | **`kubectl`이 안 된다. DNS·인그레스도 사라진다** | 클러스터 제어는 살아 있다 |
**노드 상실 실험은 agent를 죽이는 것이다.** server를 죽이는 것은 노드 상실이
아니라 **컨트롤 플레인 상실**이며 성격이 완전히 다르다.
이 사실을 모르고 "keycloak 하나만 있는 노드를 죽이자"고 계획했다가
실제 배치를 조회한 뒤 정정했다.
---
## 11층. Keycloak 클러스터링 내부 — Infinispan과 JGroups
### 두 층으로 되어 있다
```
Infinispan 분산 캐시. "세션을 어디에 두고 어떻게 복제할까"
JGroups 그룹 통신. "누가 멤버이고 어떻게 메시지를 주고받을까"
TCP 7800 실제 소켓
```
Keycloak은 Infinispan을 쓰고, Infinispan은 JGroups 위에서 돈다.
로그의 `org.infinispan.CLUSTER`와 `vendor_jgroups_*` 지표가 각각 이 두 층이다.
### 디스커버리와 트랜스포트는 다른 경로다
**이것이 이 실험대를 2노드로 만든 이유다.**
| 단계 | 경로 | 끊기면 |
|---|---|---|
| **디스커버리** — 서로를 찾는다 | PostgreSQL `JGROUPS_PING` 테이블 | 상대의 존재를 모른다 |
| **트랜스포트** — 실제로 대화한다 | **TCP 7800** | **DB엔 등록되는데 클러스터가 안 붙는다** |
`JGROUPS_PING` 한 테이블에 두 메커니즘이 다 보인다.
```
name | cluster_name | ip | coord
------------------+--------------+-----------------+-------
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
───────────────────────────── ──── ─
디스커버리 결과 트랜스포트 경로 코디네이터
```
전체 스키마는 `address / name / cluster_name / ip / coord / last_update /
coordinated_by`이고 기본키는 `address`다.
> 오래된 자료에는 `own_addr`, `ping_data` 같은 컬럼명이 나오지만 Keycloak 26의
> 실제 스키마는 위와 같다. 쿼리 전에 `\d jgroups_ping`으로 확인한다.
**`jdbc-ping`을 쓰는 이유** — 예전에는 UDP 멀티캐스트로 서로를 찾았다.
쿠버네티스나 클라우드에서는 멀티캐스트가 막혀 있는 경우가 많아,
**이미 있는 데이터베이스를 게시판처럼 쓰는** 방식으로 바뀌었다.
Keycloak 26의 기본값이다.
### 코디네이터
`coord = t` 인 노드가 **코디네이터**다. 뷰 변경을 확정하고 리밸런싱을
주도한다. 특별한 권한이 아니라 **역할**이며, 그 노드가 사라지면 남은 멤버가
인계받는다.
실험대를 전원 종료했다 켰을 때 코디네이터가 `keycloak-1` → `keycloak-0`으로
바뀌는 것을 관찰했다. **먼저 뜬 쪽이 맡는다.**
### 클러스터 뷰
```
ISPN000094: Received new cluster view for channel ISPN:
[keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938, keycloak-0-49501]
─────────────────────────── ─ ─ ────────────────────────────────────
뷰를 만든 코디네이터 뷰 ID 멤버 수 멤버 목록
```
**뷰(view)는 "지금 이 순간의 멤버 명단"** 이다. 멤버가 들어오거나 나가면
새 뷰가 발행되고 뷰 ID가 올라간다.
| 로그 코드 | 의미 |
|---|---|
| `ISPN000094` | 새 클러스터 뷰를 받았다 |
| `ISPN000079` | 자기 주소와 물리 주소(7800) |
| `ISPN100000` | 노드가 합류했다 |
```bash
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
```
### 주요 JGroups 프로토콜 — 지표 이름에 그대로 나온다
| 프로토콜 | 하는 일 | 관련 지표 |
|---|---|---|
| **GMS** (Group Membership Service) | 멤버십 관리, 뷰 발행 | `vendor_jgroups_gms_*` |
| **FD_SOCK2** (Failure Detection) | **TCP 소켓으로 상대 생존 감시** | `..._get_num_suspected_members` |
| **MERGE3** | **split brain 후 다시 합치기** | `..._merge3_get_views` |
| **NAKACK2** | 신뢰성 있는 메시지 전달, 재전송 | `..._nakack2_*` |
| **TCP** | 트랜스포트 | `..._tcp_*` |
**7800을 막으면 FD_SOCK2가 먼저 반응한다.** 소켓 연결이 끊기면 상대를
suspect 하고, GMS가 그 멤버를 뷰에서 제외한다. 각자 자기만 있는 뷰가 되면
**split brain**이고, 통신이 복구되면 MERGE3가 합친다.
### 세션은 어디에 있는가 — 두 곳이되 역할이 다르다
Keycloak 26의 기본값 `persistent-user-sessions`에서는
| 저장소 | 역할 | 노드 간 공유 |
|---|---|---|
| **PostgreSQL** | **진실의 원천.** 재시작에도 살아남는다 | **여기서만 일어난다** |
| **Infinispan `sessions`** | **자기 노드가 로그인시킨 세션만** 담는 룩어사이드 캐시 | **일어나지 않는다** |
> **처음에 이 표에 "Infinispan = 캐시 + 노드 간 실시간 전파"라고 썼는데
> 틀렸다.** 실험 0에서 측정해보니 세션 엔트리는 노드 사이를 건너가지 않는다.
> 두 노드가 같은 답을 하는 이유는 복제가 아니라 같은 DB를 보기 때문이고,
> 반대편 노드가 실제로 날리는 `SELECT ... FROM OFFLINE_USER_SESSION` 을
> PostgreSQL 로그에서 직접 잡았다.
> → [`docs/experiment-00-session-replication.md`](experiment-00-session-replication.md)
`--features-disabled=persistent-user-sessions`로 끄면 Infinispan만 남는
**volatile** 모드가 되고, 그때는 캐시가 곧 진실의 원천이므로 **복제가
반드시 일어나야 한다.** 이 둘의 차이가 로드맵 2번의 주제다.
### 세션 쓰기 트랜잭션의 세 가지 설계 결정
PostgreSQL 문장 로깅으로 잡은 갱신 트랜잭션 하나에 다 들어 있다.
| 보이는 것 | 뜻 |
|---|---|
| `update ... where ... and VERSION=$5` | **낙관적 락.** 읽을 때의 버전과 같을 때만 쓴다 |
| `for no key update ... skip locked` | 잠긴 행을 **기다리지 않고 건너뛴다.** 대기 대신 재시도 |
| **`SET LOCAL synchronous_commit TO OFF`** | **WAL 플러시를 기다리지 않고 커밋한다** |
마지막 것이 특히 중요하다 — **DB가 강제 종료되면 직전 수백 밀리초의 세션
갱신이 사라질 수 있다.** 버그가 아니라 의도된 트레이드오프다.
`LAST_SESSION_REFRESH` 갱신은 매우 잦고, 잃어도 사용자가 다시 갱신하면 된다.
---
## 12층. 관측성 — Prometheus의 구조
### 세 부분으로 되어 있다
```
수집(scrape) ──▶ 저장(TSDB) ──▶ 질의(PromQL)
15초마다 로컬 디스크 Grafana 또는 API
HTTP GET /metrics 시계열
```
**Prometheus는 pull 방식이다.** 대상이 보내주는 것이 아니라 Prometheus가
주기적으로 `/metrics`를 긁어간다.
| 결과 | |
|---|---|
| 대상이 죽으면 | 긁기가 실패하고 **`up`이 0이 된다** — 죽은 사실 자체가 데이터가 된다 |
| 방화벽 방향 | Prometheus → 대상. 대상이 Prometheus 주소를 알 필요가 없다 |
| 짧은 작업 | 긁히기 전에 끝나면 잡히지 않는다 (Pushgateway가 필요한 경우) |
### exporter 패턴
애플리케이션이 Prometheus 형식을 모를 때, **번역기**를 옆에 둔다.
| exporter | 무엇을 노출하는가 |
|---|---|
| **node-exporter** | 머신 — CPU, 메모리, 디스크, 네트워크 |
| kube-state-metrics | 쿠버네티스 오브젝트 상태 |
| postgres-exporter | PostgreSQL 내부 통계 |
**Keycloak과 Traefik은 exporter가 필요 없다.** 자체적으로 Prometheus 형식
엔드포인트를 제공한다(`KC_METRICS_ENABLED=true`).
### 서비스 디스커버리 — 타깃을 적어두지 않는다
```yaml
kubernetes_sd_configs:
- role: endpoints
namespaces: { names: [keycloak-lab] }
```
**파드 IP는 재시작마다 바뀐다.** 실험대를 전원 종료했다 켜니 모든 파드가
새 주소를 받았다(`10.42.1.22` → `10.42.1.25`). 정적 목록은 그때마다 깨진다.
`role`에 따라 무엇을 찾을지가 달라진다.
| role | 찾는 것 |
|---|---|
| `endpoints` | 서비스 뒤의 실제 파드들 ← 애플리케이션 지표 |
| `node` | 노드 |
| `pod` | 파드 직접 |
| `service` | 서비스 |
### relabel — 걸러내고 이름을 붙인다
디스커버리는 **전부 다** 가져온다. 그중 필요한 것만 남기는 것이 relabel이다.
```yaml
relabel_configs:
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
action: keep
regex: keycloak-headless;management
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
```
| `action` | 하는 일 |
|---|---|
| `keep` | regex에 맞는 것만 남긴다 |
| `drop` | 맞는 것을 버린다 |
| `replace` (기본) | 라벨 값을 만든다 |
| `labelmap` | 메타 라벨을 일반 라벨로 복사 |
**`__`로 시작하는 라벨은 내부용**이며 저장되지 않는다. `__meta_*`는
디스커버리가 붙여준 정보이고, 필요하면 `target_label`로 옮겨야 남는다.
**`pod`과 `node` 라벨을 붙이는 것이 실험에서 결정적이다.** 없으면
"어느 파드가, 어느 노드에서"에 답할 수 없다.
### 메트릭 타입
| 타입 | 성질 | 예 |
|---|---|---|
| **counter** | **누적. 줄지 않는다** (재시작 시 0으로) | `..._requests_total` |
| **gauge** | 오르내린다 | `node_memory_MemAvailable_bytes` |
| **histogram** | 구간별 분포 + 합계 + 개수 | `..._seconds_bucket/_sum/_count` |
| summary | 분위수를 클라이언트가 계산 | |
**counter는 그대로 보면 의미가 없다.** 변화율을 봐야 한다.
```promql
rate(http_requests_total[5m])
```
**histogram은 세 지표가 한 벌**이다. `_bucket`으로 분위수를 계산한다.
```promql
histogram_quantile(0.95, rate(keycloak_session_expiration_task_seconds_bucket[5m]))
```
### `up` — 가장 중요한 합성 지표
```promql
up
up{job="keycloak"}
```
Prometheus가 **직접 만드는** 지표다. 긁기에 성공하면 1, 실패하면 0.
**장애 실험에서 이것이 핵심인 이유** — 다른 지표는 대상이 죽으면 **사라진다.**
사라진 데이터로는 "언제부터 죽었나"를 알 수 없다. `up`은 **0이라는 값으로
남기 때문에** 사후에 시각을 특정할 수 있다.
```promql
up == 0 # 지금 죽은 타깃
changes(up[1h]) # 1시간 동안 몇 번 오르내렸나
min_over_time(up[10m]) # 10분 중 한 번이라도 죽었나
```
### TSDB와 보존 기간
```yaml
--storage.tsdb.path=/prometheus
--storage.tsdb.retention.time=7d
```
로컬 디스크에 시계열로 저장한다. **보존 기간이 지나면 삭제**되므로 볼륨이
무한히 커지지 않는다.
`emptyDir`에 두면 파드 재시작 시 **실험 기록이 통째로 사라진다.**
사후 추적이 목적이면 PVC여야 한다.
### 관측 시스템의 장애 도메인
**관측 시스템은 관측 대상과 같이 죽으면 안 된다.** 죽는 순간을 기록해야
하는데 같이 죽으면 기록이 없다.
노드가 둘뿐인 실험대에서는 완전히 피할 수 없으므로 **규칙으로 정한다.**
```
kc-lab-1 (server) 관측 스택을 둔다. 죽이지 않는다
kc-lab-2 (agent) 장애 주입 대상
```
`nodeSelector`로 못박아 실험이 재현 가능하게 만든다.
---
## 13층. 가상화 운영 — 실행 중 바꾸는 것들
### VM 메모리 재배분 — 게스트를 다시 만들지 않는다
```bash
virsh setmaxmem kc-lab-1 5120M --config
virsh setmem kc-lab-1 5120M --config
```
| 명령 | 바꾸는 것 |
|---|---|
| `setmaxmem` | **상한**. 부팅 시 게스트가 보는 총량 |
| `setmem` | **현재 할당**. 상한 이하여야 한다 |
**순서가 중요하다.** 현재값을 상한보다 크게 줄 수 없으므로 `setmaxmem`이
먼저다.
| 플래그 | 적용 범위 |
|---|---|
| `--config` | 영구 정의. **다음 부팅부터** |
| `--live` | 실행 중인 도메인에 즉시 |
| 둘 다 | 지금과 앞으로 |
`setmaxmem --live`는 대개 거부된다 — 게스트가 부팅 시 메모리 맵을 정하기
때문이다. **상한을 바꾸려면 게스트를 껐다 켜야 한다.**
```bash
virsh dominfo kc-lab-1 | grep -i memory
ssh kc-lab-1 free -m # 게스트가 실제로 인식한 값
```
호스트에서 8GB→12GB로 물리 증설한 뒤 이 방법으로 재배분했다.
**게스트 재생성이나 디스크 조작은 전혀 필요 없었다.**
### 안전한 종료 순서
전원을 내리기 전에 **위에서부터** 정리한다.
```bash
# 1. 애플리케이션 — 클러스터에서 정상 탈퇴
kubectl -n keycloak-lab scale statefulset/keycloak --replicas=0
kubectl -n keycloak-lab wait --for=delete pod -l app=keycloak --timeout=120s
# 2. 데이터베이스 — 마지막에, 충분한 시간을 주고
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=120s
# 3. 게스트 — ACPI 정상 종료
virsh shutdown kc-lab-1 && virsh shutdown kc-lab-2
# 4. 호스트
sudo systemctl poweroff
```
**왜 순서가 중요한가** — `virsh shutdown`은 게스트 systemd가 k3s를 멈추고,
k3s가 컨테이너에 SIGTERM을 보낸다. 유예 시간이 짧으면 **PostgreSQL이
강제 종료되어 다음 기동에 crash recovery가 돈다.** 미리 내려두면 그 위험이
없다.
**clean shutdown 확인**
```bash
ssh kc-lab-2 'sudo ls /var/lib/rancher/k3s/storage/*postgres-data*/pgdata/postmaster.pid'
```
**`postmaster.pid`가 남아 있지 않아야 정상**이다. 남아 있으면 비정상 종료였고
다음 기동에 복구 절차가 실행된다.
### 복구 순서 — 종료의 역순
```bash
virsh start kc-lab-1 && virsh start kc-lab-2
kubectl get nodes # Ready 2개 대기
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
kubectl -n keycloak-lab rollout status deployment/postgres
kubectl -n keycloak-lab scale statefulset/keycloak --replicas=2
```
**PostgreSQL이 먼저다.** Keycloak이 DB 없이 뜨면 기동에 실패한다.
**스케일을 0으로 내려두면 자동으로 복구되지 않는다.** 명시적으로 올려야 한다.
---
## 아직 기록하지 않은 개념 ## 아직 기록하지 않은 개념
실험 설계 단계에서 아래 항목을 이 문서에 추가한다. 실험을 진행하면서 이 문서에 추가한다.
- Infinispan, `DIST_SYNC`, `numOwners`, 캐시별 설정 - `persistent-user-sessions` / `volatile-user-sessions` 의 실제 차이 (로드맵 2번)
- JGroups, `JDBC_PING`, 디스커버리와 트랜스포트의 분리, TCP 7800 - refresh token rotation·revoke·max reuse 와 동시 갱신 경쟁 (로드맵 5번)
- `persistent-user-sessions` / `volatile-user-sessions`
- 원격 Infinispan(Hot Rod)과 multi-site
- refresh token rotation, revoke, max reuse, 동시 갱신 경쟁
- SSO 세션 vs 애플리케이션 세션, `KEYCLOAK_IDENTITY`, `AUTH_SESSION_ID` - SSO 세션 vs 애플리케이션 세션, `KEYCLOAK_IDENTITY`, `AUTH_SESSION_ID`
- 백채널 로그아웃과 `sid` 역인덱스 - 백채널 로그아웃과 `sid` 역인덱스
- 쿠키 `Secure` / `SameSite` / `HttpOnly`
- Redis 영속화(RDB/AOF)와 세션 복구 - Redis 영속화(RDB/AOF)와 세션 복구
- `tc netem`, OOM killer와 `oom_score`, fsync와 페이지 캐시 - Spring Session / `OAuth2AuthorizedClientService` 의 저장 구조
- `tc netem` 지연 주입
- OOM killer 와 `oom_score`
- fsync 와 페이지 캐시, EBS IOPS
### 이번에 채운 것 (2026-09-04)
10~13층으로 기록 완료 — StatefulSet·DaemonSet, PVC/PV/StorageClass,
Secret, RBAC 와 서브리소스, nodeSelector·taint, k3s server/agent 차이,
Infinispan·JGroups(디스커버리 vs 트랜스포트, GMS/FD_SOCK2/MERGE3),
Prometheus(pull·SD·relabel·메트릭 타입·`up`·TSDB), VM 메모리 재배분,
안전한 종료·복구 순서.
+400
View File
@@ -0,0 +1,400 @@
# 이 실험을 이해하기 위한 선수 지식
실험 결과를 먼저 들이밀었더니 맥락이 사라졌다. 이 문서는 **왜 이런 걸
측정하고 있는지**를 바닥부터 세운다.
읽는 순서가 곧 의존 관계다. 아는 절은 건너뛰어도 되지만, 3장까지는
"세션"이라는 말의 뜻이 계속 바뀌므로 훑고 가는 편이 낫다.
---
## 0. 출발점 — 당신이 원래 물은 것
> Keycloak이 여러 개일 경우, 세션 저장소를 Redis나 별도 저장소로 쓸 경우,
> Redis와 DB에 분리해서 세션과 토큰을 관리할 때 어떻게 달라지는지.
> SSO를 추가하면 어떻게 달라지는지. Redis 또는 DB가 죽으면 어떻게 복구하는지.
이 질문에 답하려면 **"세션이 어디에 있는가"** 를 정확히 알아야 한다.
지금 하고 있는 실험은 전부 그 한 문장을 쪼갠 것이다.
---
## 1. HTTP는 기억이 없다
모든 것의 출발점.
```
요청 1: GET /login → 서버
요청 2: GET /mypage → 서버 ← 서버는 요청 1을 기억하지 못한다
```
HTTP 요청은 **하나하나가 완전히 독립적**이다. 서버 입장에서 두 번째 요청은
생판 처음 보는 사람이 보낸 것과 구별되지 않는다.
그래서 "로그인했다"는 사실을 **어딘가 저장**해야 한다.
```
브라우저 서버
┌──────────────┐ ┌────────────────────────┐
│ 쿠키 │ │ 세션 저장소 │
│ SESSIONID= │ ──── 매 요청 ────▶ │ abc123 → { │
│ abc123 │ 이 값만 보냄 │ user: "홍길동", │
└──────────────┘ │ 로그인시각: ... │
│ } │
작은 표만 들고 다닌다 └────────────────────────┘
실제 내용은 여기 있다
```
| 용어 | 뜻 |
|---|---|
| **쿠키** | 브라우저가 들고 다니는 **작은 표(번호표)**. 보통 세션 ID만 들어 있다 |
| **세션** | 서버가 그 번호에 대해 기억하는 **실제 내용** |
**서버가 1대면 여기서 이야기가 끝난다.** 문제는 2대부터다.
---
## 2. 서버가 2대가 되는 순간 — 이 실험의 진짜 출발점
```
로그인 요청 ──▶ 서버 A A의 메모리에 "abc123 = 홍길동" 기록
다음 요청 ──▶ 서버 B B: "abc123? 그런 거 모르는데" → 로그아웃 화면
```
**이게 전부다.** 분산 세션이라는 주제 전체가 이 한 장면에서 나온다.
푸는 방법은 셋뿐이다.
| 방법 | 어떻게 | 대가 |
|---|---|---|
| **1. 고정 배정** (sticky session) | 같은 사람은 항상 같은 서버로 보낸다 | **그 서버가 죽으면 그 사람 세션은 사라진다.** 부하도 안 고르게 퍼진다 |
| **2. 복제** | 서버끼리 메모리 내용을 서로 보낸다 | 서버가 N대면 트래픽이 N² 로 는다. 어긋남(불일치)이 생긴다 |
| **3. 공유 저장소** | 세션을 바깥(DB·Redis)에 두고 모두가 본다 | **그게 죽으면 전체가 멈춘다.** 매 요청마다 네트워크 왕복 |
> **당신이 원래 물은 "Redis나 별도 저장소를 쓰면"이 바로 3번**이다.
> 그리고 "Redis나 DB가 죽으면 어떻게 복구하나"는 3번의 대가를 묻는 것이다.
**Keycloak도 예외가 아니다.** Keycloak을 2대 띄우면 정확히 이 문제가 생긴다.
Keycloak이 이걸 어떻게 풀었는지가 실험 0의 주제다.
---
## 3. Keycloak은 무엇이고, 왜 세션을 갖는가
### 3-1. 하는 일
Keycloak은 **로그인을 대신 해주는 서버**다.
```
[사용자] [내 앱] [Keycloak]
│ │ │
│─ 접속 ──────▶│ │
│◀─ "Keycloak 가서 로그인하고 와" ────│
│──────────────────── 로그인 ───────▶│
│◀─────────────── 토큰 발급 ─────────│
│─ 토큰 들고 ──▶│ │
│ │─ 이 토큰 유효해? ──▶│
```
내 앱은 비밀번호를 저장하지도, 검증하지도 않는다. 그 일을 Keycloak이 한다.
### 3-2. 그래서 **세션이 두 겹**이 된다
여기가 헷갈리는 지점이다. "세션"이라는 말이 두 가지를 가리킨다.
```
┌─────────────────────────────────────────────────────┐
│ Keycloak 의 SSO 세션 │
│ "이 브라우저는 홍길동으로 로그인되어 있다" │
│ 쿠키 이름: KEYCLOAK_IDENTITY │
└─────────────────────────────────────────────────────┘
│ │
▼ ▼
┌──────────────────┐ ┌──────────────────┐
│ 앱1 의 세션 │ │ 앱2 의 세션 │
│ (또는 토큰) │ │ (또는 토큰) │
└──────────────────┘ └──────────────────┘
```
| | 누가 갖는가 | 사라지면 |
|---|---|---|
| **SSO 세션** | **Keycloak** | 모든 앱에서 다시 로그인해야 한다 |
| 앱 세션 | 각 애플리케이션 | 그 앱만 다시 들어가면 된다 |
**SSO가 되는 원리가 이것이다.** 앱1에서 로그인하면 Keycloak에 SSO 세션이
생긴다. 앱2로 가면 Keycloak이 "이 브라우저 이미 로그인했네" 하고 **로그인
화면 없이** 바로 토큰을 준다.
> **그래서 Keycloak의 세션이 사라지면 SSO 전체가 깨진다.**
> 당신 질문의 "SSO를 추가하면 어떻게 달라지는지"가 여기 걸린다.
> 앱이 하나일 때는 그 앱만 재로그인이지만, SSO에서는 **전 앱이 동시에** 터진다.
---
## 4. 토큰이 있는데 왜 세션이 필요한가
가장 흔한 오해다. "JWT는 stateless라서 서버가 기억할 게 없다"는 말은
**반만 맞다.**
### 4-1. 토큰이 두 종류다
```
로그인 성공
├──▶ access token 수명 짧음 (이 실험대: 60초)
│ JWT. 서명이 붙어 있어 서버가 아무것도 기억 안 해도 검증된다
│ → 진짜 stateless
└──▶ refresh token 수명 김 (이 실험대: 1800초 = 30분)
access token 이 만료되면 이걸로 새로 받는다
→ 서버가 세션을 기억하고 있어야 한다
```
| | access token | refresh token |
|---|---|---|
| 검증 방식 | **서명만 보면 됨** | **서버 세션 조회 필요** |
| 취소 | **불가능** (만료를 기다려야) | 가능 |
| 수명 | 짧게 (분 단위) | 길게 (시간~일) |
### 4-2. 그래서 이렇게 된다
```
0초 로그인 세션 생성
0초 access token 발급 이후 60초간은 서버에 안 물어봐도 됨
60초 access token 만료
60초 refresh 요청 ─────▶ 서버: "이 세션 살아 있나?" ◀── 여기서 세션 필요
60초 새 access token
120초 또 만료 → 또 refresh → 또 세션 조회
```
**60초마다 세션 저장소를 친다.** access token 수명이 짧을수록 세션 저장소
부하가 커진다 — 보안과 성능의 맞바꿈이 여기서 일어난다.
### 4-3. 로그아웃도 세션이 있어야 한다
**로그아웃 = 세션 삭제**다. 세션이 없으면 로그아웃이라는 개념 자체가 없다.
이미 발급된 access token은 서명이 유효하므로 만료 전까지 계속 통과한다.
> 그래서 access token 수명을 60초로 짧게 잡는다. 로그아웃해도 최대 60초는
> 살아 있다는 뜻이고, 그 이상은 refresh 가 막히므로 끝난다.
---
## 5. Keycloak은 세션을 어디에 두는가 — **버전에 따라 답이 다르다**
이 실험 전체가 여기에 걸려 있다.
### 5-1. 두 개의 후보
| | 무엇 | 성질 |
|---|---|---|
| **Infinispan** | Keycloak **안에 내장된** 분산 캐시. Java 라이브러리 | 메모리. 빠름. 프로세스가 죽으면 사라짐 |
| **데이터베이스** | PostgreSQL 등 바깥의 DB | 디스크. 느림. 재시작해도 남음 |
**Infinispan은 별도로 설치하는 물건이 아니다.** Keycloak 프로세스 안에서 도는
라이브러리다. Redis처럼 따로 띄우는 게 아니다 — 이걸 헷갈리면 전체가 안 맞는다.
### 5-2. 버전별로 이렇게 바뀌었다
| 버전 | 진실의 원천 | 전체 재시작하면 |
|---|---|---|
| ~24 | **Infinispan (메모리)** | **세션 전부 소멸** |
| 25 | 선택 (`persistent-user-sessions` 옵션) | 설정에 따라 |
| **26 (지금 이 실험대)** | **데이터베이스** | **세션 살아남음** |
**이게 결정적이다.** 인터넷에 있는 Keycloak 클러스터링 자료 대부분은
24 이전 기준이라 **"세션은 Infinispan이 노드끼리 복제한다"** 고 쓰여 있다.
26에서는 더 이상 사실이 아니다.
> 제가 처음에 개념 문서에 "Infinispan = 캐시 + 노드 간 실시간 전파"라고
> 써둔 것도 이 옛 모델을 그대로 옮긴 것이었다. 실험 0에서 틀렸음이 드러났다.
---
## 6. Infinispan / JGroups / 7800 — 이름들의 정체
실험 로그에 계속 나오는 이름들이다.
```
Keycloak 프로세스
┌────────────────────────────────────────┐
│ Infinispan "세션을 어디 두고 어떻게 │ ← 캐시 계층
│ 나눌까" │
│ │ │
│ JGroups "누가 우리 멤버이고 │ ← 그룹 통신 계층
│ 어떻게 메시지를 주고받나" │
│ │ │
│ TCP 7800 실제 소켓 │ ← 네트워크
└────────────────────────────────────────┘
```
| 이름 | 정체 |
|---|---|
| **Infinispan** | Keycloak 내장 캐시. 세션·realm 설정·로그인 실패 횟수 등을 담는다 |
| **JGroups** | Infinispan이 노드끼리 대화할 때 쓰는 하부 라이브러리 |
| **TCP 7800** | JGroups가 쓰는 포트. **노드 간 통신 경로** |
| **jdbc-ping** | 서로를 **찾는** 방법. DB의 `JGROUPS_PING` 테이블을 게시판처럼 쓴다 |
| `ISPN000094` | "새 멤버 명단을 받았다"는 로그 코드 |
**찾는 것과 대화하는 것이 다른 경로다.**
```
디스커버리 (서로를 찾는다) → PostgreSQL JGROUPS_PING 테이블
트랜스포트 (실제 대화) → TCP 7800
```
---
## 7. 왜 쿠버네티스와 노드 2대가 나오는가
당신 질문은 "Keycloak이 여러 개일 경우"였다. 그걸 **진짜로** 재현하려면
Keycloak 프로세스 2개가 **서로 다른 기계**에 있어야 한다.
| 방식 | 노드 상실을 실험할 수 있나 |
|---|---|
| Docker 컨테이너 2개 (한 기계) | **못 한다.** 커널이 하나라 "기계가 죽는" 상황을 못 만든다 |
| **VM 2대 + k3s** | **된다.** 하나를 전원 차단할 수 있다 |
그래서 이 실험대는 VM 2대(`kc-lab-1`, `kc-lab-2`) 위에 k3s를 올렸다.
```
kc-lab-1 (k3s 서버) kc-lab-2 (k3s 에이전트)
├─ keycloak-1 ├─ keycloak-0
├─ traefik, coredns └─ postgres
└─ prometheus, grafana
```
**`keycloak-0` / `keycloak-1` 은 Keycloak 프로세스**이고,
**`kc-lab-1` / `kc-lab-2` 는 그것들이 올라간 기계**다. 이름이 비슷해서
헷갈리기 쉬운데 계층이 다르다.
---
## 8. 그래서 실험 0은 무엇을 알아내려 한 것인가
### 8-1. 답해야 할 실무 질문
> Keycloak을 2대로 늘렸다. **한 대가 죽으면 로그인한 사람들은 어떻게 되나?**
> **DB가 죽으면?** **노드 사이 네트워크가 끊기면?**
이 질문들에 답하려면 **정상일 때 무엇이 어디에 있는지**를 먼저 알아야 한다.
그게 없으면 장애를 일으켜도 무엇이 왜 깨졌는지 해석할 수 없다.
### 8-2. 그래서 실험 0의 질문은 두 개다
```
질문 A. 한 노드에서 만든 세션을 다른 노드가 쓸 수 있는가?
↓ 답: 그렇다
질문 B. 그 공유는 무엇 덕분인가?
(a) Infinispan 이 메모리를 복제해서
(b) 둘 다 같은 DB 를 봐서
```
### 8-3. **B를 구분해야 하는 이유** — 운영 대응이 정반대다
| 상황 | (a) 복제라면 | (b) DB라면 |
|---|---|---|
| 노드 간 7800 끊김 | **세션 공유 깨짐** | **멀쩡** |
| DB 죽음 | 한동안 버팀 | **즉시 전면 장애** |
| 노드 1대 죽음 | 세션 살아남음 | 세션 살아남음 |
| 성능 병목 | 노드 간 네트워크 | **DB, 커넥션 풀** |
| 튜닝할 곳 | JGroups 설정 | **DB 인덱스, 커넥션 수** |
| 노드를 10대로 늘리면 | **복제 트래픽 폭증** | DB 부하 증가 |
**같은 증상에 정반대 처방이 나온다.** 그래서 추측이 아니라 측정으로
확정해야 했다.
---
## 9. 왜 하필 "캐시 엔트리 개수"를 셌는가
질문 B를 가르는 가장 직접적인 방법이기 때문이다.
```
keycloak-0 에만 로그인을 보낸다
└──▶ 그리고 keycloak-1 의 메모리를 들여다본다
그 세션이 들어와 있으면 → (a) 복제한 것
비어 있으면 → (b) DB 로 공유한 것
```
Keycloak은 자기 캐시에 몇 개가 들었는지를 `/metrics` 로 알려준다.
```
vendor_statistics_approximate_entries_unique{cache="sessions"} 7.0
───────────── ───
세션 캐시 7개 들어 있다
```
**측정 결과: keycloak-1은 계속 0이었다.** keycloak-0이 14개를 들고 있는
동안에도 0. 그리고 keycloak-1에 직접 로그인을 보낸 순간에만 늘었다.
```
단계 k0 k1
시작 2 0
keycloak-1 에 로그인 5회 2 5 ← k0 안 늘어남
keycloak-0 에 로그인 5회 7 5 ← k1 안 늘어남
PostgreSQL 세션 수: 12 = 7 + 5 ← 캐시 합과 정확히 일치
```
**각 노드는 자기가 처리한 것만 캐시한다. 메모리는 건너가지 않는다.**
→ 답은 **(b)**.
그리고 마지막으로 PostgreSQL 로그를 켜서, keycloak-1이 **실제로 날리는
SELECT 문**을 잡았다. 추측이 아니라는 것을 못 박기 위해서다.
---
## 10. 이 사실이 당신 운영에 뜻하는 것
| 알게 된 것 | 실무적 의미 |
|---|---|
| 세션은 DB에 있다 | **DB가 단일 장애점이다.** HA·백업 계획이 Keycloak 대수보다 중요하다 |
| 메모리는 로컬 캐시일 뿐 | Keycloak을 몇 대로 늘려도 **노드 간 트래픽은 안 는다.** 대신 DB 부하가 는다 |
| 남의 세션은 캐시 안 함 | **sticky session 은 정확성이 아니라 성능 문제다.** 없어도 동작하지만 DB를 더 친다 |
| refresh 마다 DB 읽기+쓰기 | access token 수명을 줄이면 **DB 부하가 그만큼 는다** |
| `synchronous_commit OFF` | **DB가 강제 종료되면 직전 수백 ms 갱신이 사라진다** (의도된 설계) |
| 낙관적 락 (`VERSION`) | **동시에 refresh 하면 한쪽이 진다.** 클라이언트에 재시도가 필요하다 |
---
## 11. 앞으로 할 실험과 각각이 답하는 질문
| # | 실험 | 답하는 실무 질문 | 예측 |
|---|---|---|---|
| **A-0** | ✅ 세션 복제 확인 | 정상일 때 세션은 어디 있나 | — (완료) |
| **A-1** | TCP 7800 차단 | 노드 간 네트워크가 끊기면? | 세션 공유는 **안 깨짐**. 무효화 전파가 깨질 것 |
| **A-2** | DB 정지 | **DB가 죽으면?** | **즉시 전면 장애** |
| A-2' | DB 강제 종료 | 복구하면 뭘 잃나 | 직전 수백 ms 세션 갱신 소멸 |
| **A-3** | 노드 1대 전원 차단 | **Keycloak 한 대가 죽으면?** | 세션 살아남음 |
| **A-4** | volatile 모드 비교 | 옛 방식(24 이전)은 뭐가 다른가 | A-1이 **정반대로** 치명적이 됨 |
| B-5 | 동시 refresh 경쟁 | 토큰 갱신이 겹치면? | 한쪽이 낙관적 락에서 짐 |
| — | SSO 다중 앱 | **SSO를 붙이면 뭐가 달라지나** | Keycloak 세션 하나가 전 앱을 좌우 |
**A-1이 특히 중요하다.** 통념("클러스터 포트 막으면 세션 깨짐")과
이번 측정("세션은 7800으로 안 다님")이 정면으로 어긋나므로, **둘 중 하나는
틀렸다.** 실험이 판정한다.
---
## 12. 용어 빠른 참조
| 용어 | 한 줄 |
|---|---|
| **세션** | 서버가 "이 사람 로그인했음"을 기억하는 것 |
| **SSO 세션** | Keycloak이 가진 세션. 이게 죽으면 전 앱 재로그인 |
| **access token** | 짧게 사는 JWT. 서명만으로 검증. 취소 불가 |
| **refresh token** | access token을 새로 받는 표. **세션 조회가 필요** |
| **sid** | 세션 식별자. JWT·DB·관리 API에서 **같은 문자열** |
| **Infinispan** | Keycloak **내장** 캐시 (별도 설치 아님) |
| **JGroups** | Infinispan의 노드 간 통신 라이브러리 |
| **jdbc-ping** | DB 테이블로 서로를 찾는 방식 |
| **7800** | 노드 간 통신 포트 |
| **persistent-user-sessions** | 세션을 DB에 저장하는 기능. **KC 26 기본값** |
| **`OFFLINE_USER_SESSION`** | 이름과 달리 **온라인 세션도** 여기 있다 (`offline_flag='0'`) |
| **낙관적 락** | 읽을 때 버전과 같을 때만 쓰기. 충돌은 사후 검출 |
| **kc-lab-1/2** | VM(기계) 이름 |
| **keycloak-0/1** | Keycloak 프로세스(파드) 이름 |
+6 -1
View File
@@ -59,7 +59,9 @@ if [ "$expected_count" -ne 39 ]; then
12. 인증서 갱신 실측 12. 인증서 갱신 실측
``` ```
**순서 근거는 [`open-questions-coverage.md`](open-questions-coverage.md)에 있다.** **각 실험의 구조·주입 지점·확인 항목은
[`experiment-plan.md`](experiment-plan.md)에 미리 확정해두었다.**
순서 근거는 [`open-questions-coverage.md`](open-questions-coverage.md)에 있다.
특히 5번(refresh 경쟁)은 3번(저장소 공유) 이후여야 **재현 자체가 성립한다.** 특히 5번(refresh 경쟁)은 3번(저장소 공유) 이후여야 **재현 자체가 성립한다.**
### ✅ 완료 — 환경 구축 ### ✅ 완료 — 환경 구축
@@ -378,6 +380,9 @@ sudo certbot renew --force-renewal
| 문서 | 내용 | | 문서 | 내용 |
|---|---| |---|---|
| [`experiment-plan.md`](experiment-plan.md) | **실험 20개의 구조도·주입 방법·관측 지점·예측** |
| [`session-lab-prerequisites.md`](session-lab-prerequisites.md) | 이 실험을 이해하기 위한 선수 지식 |
| [`experiment-00-session-replication.md`](experiment-00-session-replication.md) | A-0 결과 — 세션 공유의 실체 |
| [`open-questions-coverage.md`](open-questions-coverage.md) | 공개 열린 질문 4개와의 대조, 순서 근거 | | [`open-questions-coverage.md`](open-questions-coverage.md) | 공개 열린 질문 4개와의 대조, 순서 근거 |
| [`session-lab-concepts.md`](session-lab-concepts.md) | 등장 개념 전체 (가상화·네트워크·k3s·TLS·패키지) | | [`session-lab-concepts.md`](session-lab-concepts.md) | 등장 개념 전체 (가상화·네트워크·k3s·TLS·패키지) |
| [`session-lab-operations.md`](session-lab-operations.md) | 관측 도구 · 자주 쓰는 명령 · 훈련 · 자원 예산 | | [`session-lab-operations.md`](session-lab-operations.md) | 관측 도구 · 자주 쓰는 명령 · 훈련 · 자원 예산 |