Compare commits
14
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d0666c5ba0 | ||
|
|
99b689e715 | ||
|
|
4177fb6a48 | ||
|
|
2a98ef1090 | ||
|
|
9c3cde457e | ||
|
|
f1e8c35805 | ||
|
|
1de6108157 | ||
|
|
22d873eb4f | ||
|
|
e5ebaeb623 | ||
|
|
006da7d490 | ||
|
|
d5cc2b55a9 | ||
|
|
0bb0e0ac49 | ||
|
|
33878e8880 | ||
|
|
6dce35ec83 |
@@ -0,0 +1,22 @@
|
||||
[ 1289ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/login:0
|
||||
[ 1417ms] [VERBOSE] [DOM] Input elements should have autocomplete attributes (suggested: "username"): (More info: https://goo.gl/9p2vKq) %o @ https://app2.hyeonworks.com/login:0
|
||||
[ 9024ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 9130ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/:0
|
||||
[ 9456ms] [WARNING] Deprecation warning: value provided is not in a recognized RFC2822 or ISO format. moment construction falls back to js Date(), which is not reliable across all browsers and versions. Non RFC2822/ISO date formats are discouraged. Please refer to http://momentjs.com/guides/#/warnings/js-date/ for more info.
|
||||
Arguments:
|
||||
[0] _isAMomentObject: true, _isUTC: false, _useUTC: false, _l: undefined, _i: Thu, 27 Aug 2026 13:03:49, _f: undefined, _strict: undefined, _locale: [object Object]
|
||||
Error
|
||||
at a.createFromInputFallback (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:3)
|
||||
at an (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:25647)
|
||||
at un (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29355)
|
||||
at aa (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29221)
|
||||
at on (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:28938)
|
||||
at sa (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29715)
|
||||
at A (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:624:29748)
|
||||
at a (https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:621:89)
|
||||
at f (https://app2.hyeonworks.com/public/build/3719.c065b2e146c4c8347d51.js:1:4635)
|
||||
at u (https://app2.hyeonworks.com/public/build/322.177b4bb01c5d74f9b28f.js:2473:47448) @ https://app2.hyeonworks.com/public/build/6029.0549a3fcb50e73c4b256.js:620
|
||||
[ 9605ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 10620ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 11527ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 13875ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -0,0 +1,7 @@
|
||||
[ 144ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 153ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore:0
|
||||
[ 1077ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 2101ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 2922ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 7323ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 9370ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -0,0 +1,8 @@
|
||||
[ 271ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&orgId=1&panes=%7B%22a%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22vendor_statistics_approximate_entries_unique%7Bcache%3D%5C%22sessions%5C%22%7D%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22%7B%7Bpod%7D%7D%20on%20%7B%7Bnode%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-15m%22%2C%22to%22%3A%22now%22%7D%7D%7D:0
|
||||
[ 346ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1512ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 2433ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 6941ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 13188ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 21578ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 25998ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -0,0 +1,7 @@
|
||||
[ 766ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step1.html:0
|
||||
[ 781ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step2.html:0
|
||||
[ 17929ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step1.html:0
|
||||
[ 17949ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/3p-cookies/step2.html:0
|
||||
[ 17981ms] [WARNING] An iframe which has both allow-scripts and allow-same-origin for its sandbox attribute can escape its sandboxing. @ https://auth.hyeonworks.com/realms/master/protocol/openid-connect/login-status-iframe.html:0
|
||||
[ 18447ms] [WARNING] For accessibility reasons an aria-label should be specified on nav groups if a title isn't @ https://auth.hyeonworks.com/resources/9v5yc/admin/keycloak.v2/assets/main-BbID33M6.js:7
|
||||
[ 18462ms] [WARNING] For accessibility reasons an aria-label should be specified on nav groups if a title isn't @ https://auth.hyeonworks.com/resources/9v5yc/admin/keycloak.v2/assets/main-BbID33M6.js:7
|
||||
@@ -0,0 +1,2 @@
|
||||
[ 75ms] [ERROR] Failed to load resource: the server responded with a status of 404 (Not Found) @ https://hyeonworks.com/questions:0
|
||||
[ 100ms] [ERROR] Failed to load resource: the server responded with a status of 404 (Not Found) @ https://hyeonworks.com/favicon.ico:0
|
||||
+1
-1
@@ -1 +1 @@
|
||||
[ 295ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
[ 149ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
+1
-1
@@ -1 +1 @@
|
||||
[ 125ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
[ 103ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
+1
-1
@@ -1 +1 @@
|
||||
[ 91ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
[ 132ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
+1
-1
@@ -1 +1 @@
|
||||
[ 88ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
[ 239ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
@@ -0,0 +1 @@
|
||||
[ 105ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
@@ -0,0 +1 @@
|
||||
[ 108ms] [ERROR] Failed to load resource: the server responded with a status of 401 (Unauthorized) @ https://hyeonworks.com/api/v1/studio/session:0
|
||||
@@ -0,0 +1,41 @@
|
||||
[ 344ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 869ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&panes=%7B%22cf1%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22vendor_cluster_size%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22%7B%7Bpod%7D%7D+on+%7B%7Bnode%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-45m%22%2C%22to%22%3A%22now%22%7D%7D%7D&orgId=1:0
|
||||
[ 1014ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 2039ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 4085ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 7289ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 15556ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 26818ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 32659ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 35930ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 53849ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 61722ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 79450ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 84571ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 87233ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 92665ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 113244ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 119181ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 137390ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 157071ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 166091ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 182011ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 188613ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 206228ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 218561ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 229607ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 235818ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 237158ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 252095ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 261118ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 273809ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 280227ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 292650ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 306477ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 309957ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 311984ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 325683ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 344586ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 357576ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 359294ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 364740ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -0,0 +1,129 @@
|
||||
[ 615ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&panes=%7B%22te0%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22up%7Bjob%3D%5C%22keycloak%5C%22%7D%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22up+%E2%80%94+%7B%7Bpod%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-20m%22%2C%22to%22%3A%22now%22%7D%7D%7D&orgId=1:0
|
||||
[ 929ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 2566ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 5541ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 7990ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 12402ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 17640ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 20285ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 28277ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 34341ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 34985ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 50081ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 65649ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 75046ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 75890ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 96067ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 113467ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 121972ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 139681ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 150971ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 161211ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 180744ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 191401ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 210668ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 229082ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 239345ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 252119ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 266045ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 276705ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 289911ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 296875ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 315099ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 325431ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 328418ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 343774ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 347054ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 352405ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 361291ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 370639ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 374286ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 377413ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 394362ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 405423ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 425476ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 434299ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 441775ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 450786ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 456829ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 468812ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 478913ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 491241ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 495228ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 502295ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 511625ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 524619ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 533050ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 553189ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 571699ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 575413ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 588005ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 599161ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 608816ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 619031ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 626612ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 639003ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 655077ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 674532ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 677197ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 696027ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 709512ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 727505ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 747139ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 754565ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 773682ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 784674ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 787399ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 802845ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 825589ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 832768ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 842043ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 852717ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 871748ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 876571ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 894491ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 895515ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 904310ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 916681ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 934752ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 941175ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 946004ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 960849ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 962180ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 984446ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 997579ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1010105ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1027190ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1043388ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1061311ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1066040ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1080349ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1092336ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1104312ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1109362ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1127047ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1134728ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1154899ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1161253ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1174971ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1184491ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1185340ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1189111ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1196984ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1204794ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1247961ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1286322ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1319329ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1362771ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1396881ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1434184ms] [ERROR] Failed to load resource: the server responded with a status of 502 () @ https://app2.hyeonworks.com/api/user/auth-tokens/rotate:0
|
||||
[ 1443506ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1465016ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 502 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1513337ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1556906ms] [WARNING] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: WebSocket is closed before the connection is established. @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1562963ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 502 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1572896ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 502 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1576390ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 503 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1593794ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: HTTP Authentication failed; no valid credentials available @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1600651ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: HTTP Authentication failed; no valid credentials available @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 1616424ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: HTTP Authentication failed; no valid credentials available @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -0,0 +1,20 @@
|
||||
[ 268ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/login:0
|
||||
[ 379ms] [VERBOSE] [DOM] Input elements should have autocomplete attributes (suggested: "username"): (More info: https://goo.gl/9p2vKq) %o @ https://app2.hyeonworks.com/login:0
|
||||
[ 10453ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 10532ms] [WARNING] <meta name="apple-mobile-web-app-capable" content="yes"> is deprecated. Please include <meta name="mobile-web-app-capable" content="yes"> @ https://app2.hyeonworks.com/explore?schemaVersion=1&panes=%7B%22ich%22%3A%7B%22datasource%22%3A%22PBFA97CFB590B2093%22%2C%22queries%22%3A%5B%7B%22refId%22%3A%22A%22%2C%22expr%22%3A%22up%7Bjob%3D%7E%5C%22keycloak%7Cnode-exporter%5C%22%7D%22%2C%22range%22%3Atrue%2C%22instant%22%3Afalse%2C%22editorMode%22%3A%22code%22%2C%22legendFormat%22%3A%22%7B%7Bjob%7D%7D+%E2%80%94+%7B%7Bpod%7D%7D%7B%7Bnode%7D%7D%22%2C%22datasource%22%3A%7B%22type%22%3A%22prometheus%22%2C%22uid%22%3A%22PBFA97CFB590B2093%22%7D%7D%5D%2C%22range%22%3A%7B%22from%22%3A%22now-55m%22%2C%22to%22%3A%22now%22%7D%7D%7D&orgId=1:0
|
||||
[ 11687ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 13737ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 15113ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 20185ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 27259ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 40154ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 41870ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 50493ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 54082ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 60121ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 74362ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 83269ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 95349ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 98008ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 104458ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
[ 123816ms] [ERROR] WebSocket connection to 'wss://app2.hyeonworks.com/api/live/ws' failed: Error during WebSocket handshake: Unexpected response code: 400 @ https://app2.hyeonworks.com/public/build/1518.a3f1f690c084a37f01c7.js:362
|
||||
@@ -1,25 +0,0 @@
|
||||
- generic [ref=f7e3]:
|
||||
- link "본문으로 건너뛰기" [ref=f7e4] [cursor=pointer]:
|
||||
- /url: "#main-content"
|
||||
- banner [ref=f7e5]:
|
||||
- generic [ref=f7e6]:
|
||||
- link "TechLog 홈" [ref=f7e8] [cursor=pointer]:
|
||||
- /url: /
|
||||
- text: TechLog
|
||||
- generic [ref=f7e9]:
|
||||
- button "TechLog 검색 열기" [ref=f7e11] [cursor=pointer]: 검색
|
||||
- group [ref=f7e12]:
|
||||
- generic "메뉴" [ref=f7e13] [cursor=pointer]
|
||||
- generic [ref=f7e14]:
|
||||
- paragraph [ref=f7e15]: 화면을 준비하고 있습니다.
|
||||
- generic [ref=f7e16]: TechLog 로딩 중
|
||||
- contentinfo [ref=f7e17]:
|
||||
- generic [ref=f7e18]:
|
||||
- generic [ref=f7e19]:
|
||||
- paragraph [ref=f7e20]: 동현
|
||||
- paragraph [ref=f7e21]: 문제를 재현하고 검증해 실제 운영에 적용할 수 있는 형태로 정리합니다.
|
||||
- generic [ref=f7e22]:
|
||||
- link "프로필" [ref=f7e23] [cursor=pointer]:
|
||||
- /url: /profile
|
||||
- link "변경 기록" [ref=f7e24] [cursor=pointer]:
|
||||
- /url: /releases
|
||||
@@ -0,0 +1,2 @@
|
||||
- main [ref=e7]:
|
||||
- status "Loading" [ref=e10]
|
||||
@@ -0,0 +1 @@
|
||||
- main [ref=f3e7]
|
||||
@@ -0,0 +1,2 @@
|
||||
- main [ref=f6e7]:
|
||||
- status "Loading" [ref=f6e10]
|
||||
@@ -0,0 +1,104 @@
|
||||
- generic [ref=f9e4]:
|
||||
- link "Skip to main content" [ref=f9e5] [cursor=pointer]:
|
||||
- /url: "#pageContent"
|
||||
- banner [ref=f9e7]:
|
||||
- generic [ref=f9e8]:
|
||||
- link [ref=f9e10] [cursor=pointer]:
|
||||
- /url: /
|
||||
- img "Grafana" [ref=f9e11]
|
||||
- generic [ref=f9e14]:
|
||||
- button "Search or jump to..." [ref=f9e18] [cursor=pointer]
|
||||
- generic [ref=f9e19]: ctrl+k
|
||||
- generic [ref=f9e23]:
|
||||
- button "New" [ref=f9e24] [cursor=pointer]
|
||||
- button "Help" [ref=f9e30] [cursor=pointer]
|
||||
- button "News" [ref=f9e33] [cursor=pointer]
|
||||
- button "Profile" [ref=f9e36] [cursor=pointer]:
|
||||
- img "User avatar" [ref=f9e37]
|
||||
- generic [ref=f9e38]:
|
||||
- button "Open menu" [ref=f9e40] [cursor=pointer]
|
||||
- navigation "Breadcrumbs" [ref=f9e43]:
|
||||
- list [ref=f9e44]:
|
||||
- listitem [ref=f9e45]:
|
||||
- link "Home" [ref=f9e46] [cursor=pointer]:
|
||||
- /url: /
|
||||
- listitem [ref=f9e50]:
|
||||
- link "Explore" [ref=f9e51] [cursor=pointer]:
|
||||
- /url: /explore
|
||||
- listitem [ref=f9e55]:
|
||||
- generic "Prometheus" [ref=f9e56]
|
||||
- generic [ref=f9e57]:
|
||||
- button "Show more items" [ref=f9e60] [cursor=pointer]
|
||||
- button "Toggle top search bar" [ref=f9e64] [cursor=pointer]
|
||||
- main [ref=f9e70]:
|
||||
- generic [ref=f9e72]:
|
||||
- heading "Explore" [level=1] [ref=f9e73]
|
||||
- generic [ref=f9e78]:
|
||||
- navigation "Explore toolbar" [ref=f9e80]:
|
||||
- navigation "Search links" [ref=f9e82]:
|
||||
- generic [ref=f9e83]:
|
||||
- button "Content outline" [expanded] [ref=f9e85] [cursor=pointer]:
|
||||
- generic [ref=f9e88]: Outline
|
||||
- generic [ref=f9e93] [cursor=pointer]:
|
||||
- img "Prometheus logo" [ref=f9e95]
|
||||
- textbox "Select a data source" [ref=f9e96]:
|
||||
- /placeholder: ""
|
||||
- button "Show more items" [ref=f9e102] [cursor=pointer]
|
||||
- generic [ref=f9e106]:
|
||||
- generic [ref=f9e110]:
|
||||
- button "Collapse outline" [expanded] [ref=f9e112] [cursor=pointer]:
|
||||
- img "arrow-from-right" [ref=f9e113]
|
||||
- button "Queries" [ref=f9e116] [cursor=pointer]:
|
||||
- img "arrow" [ref=f9e117]
|
||||
- generic [ref=f9e124]:
|
||||
- generic [ref=f9e126]:
|
||||
- generic "Query editor row" [ref=f9e129]:
|
||||
- generic [ref=f9e130]:
|
||||
- generic [ref=f9e132]:
|
||||
- generic [ref=f9e133]:
|
||||
- button "Collapse query row" [expanded] [ref=f9e134] [cursor=pointer]
|
||||
- generic [ref=f9e137]:
|
||||
- button "Query editor row title A" [ref=f9e138] [cursor=pointer]:
|
||||
- generic [ref=f9e139]: A
|
||||
- emphasis [ref=f9e140]: (Prometheus)
|
||||
- generic [ref=f9e141]:
|
||||
- button "Show data source help" [ref=f9e143] [cursor=pointer]
|
||||
- button "Duplicate query" [ref=f9e147] [cursor=pointer]
|
||||
- button "Hide response" [ref=f9e151] [cursor=pointer]
|
||||
- button "Remove query" [ref=f9e155] [cursor=pointer]
|
||||
- button "Drag and drop to reorder" [ref=f9e158]:
|
||||
- img "Drag and drop to reorder" [ref=f9e159]
|
||||
- generic [ref=f9e162]:
|
||||
- generic [ref=f9e163]:
|
||||
- button "Kick start your query" [ref=f9e164] [cursor=pointer]
|
||||
- generic [ref=f9e167]:
|
||||
- generic [ref=f9e168] [cursor=pointer]: Explain
|
||||
- generic [ref=f9e169]:
|
||||
- checkbox "Explain Toggle switch" [ref=f9e170]
|
||||
- generic "Toggle switch" [ref=f9e171] [cursor=pointer]
|
||||
- radiogroup [ref=f9e176]:
|
||||
- generic [ref=f9e177]:
|
||||
- radio "Builder" [ref=f9e178] [cursor=pointer]
|
||||
- generic [ref=f9e179] [cursor=pointer]: Builder
|
||||
- generic [ref=f9e180]:
|
||||
- radio "Code" [checked] [ref=f9e181] [cursor=pointer]
|
||||
- generic [ref=f9e182] [cursor=pointer]: Code
|
||||
- generic [ref=f9e184]:
|
||||
- generic [ref=f9e186]:
|
||||
- button "Loading metrics..." [disabled] [ref=f9e187] [cursor=pointer]
|
||||
- generic [ref=f9e190]: Loading editor
|
||||
- 'button "Options Legend: {{pod}} on {{node}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=f9e198] [cursor=pointer]':
|
||||
- generic [ref=f9e202]:
|
||||
- heading "Options" [level=6] [ref=f9e203]
|
||||
- generic [ref=f9e204]:
|
||||
- generic [ref=f9e205]: "Legend: {{pod}} on {{node}}"
|
||||
- generic [ref=f9e206]: "Format: Time series"
|
||||
- generic [ref=f9e207]: "Step: auto"
|
||||
- generic [ref=f9e208]: "Type: Range"
|
||||
- generic [ref=f9e209]: "Exemplars: false"
|
||||
- generic [ref=f9e210]:
|
||||
- button "Add query" [ref=f9e211] [cursor=pointer]
|
||||
- button "Query history" [ref=f9e215] [cursor=pointer]
|
||||
- button "Query inspector" [ref=f9e219] [cursor=pointer]
|
||||
- generic:
|
||||
- main
|
||||
@@ -0,0 +1,4 @@
|
||||
- main [ref=f12e3]:
|
||||
- generic [ref=f12e4]:
|
||||
- progressbar "Contents" [ref=f12e5]
|
||||
- paragraph [ref=f12e8]: Loading the Administration Console
|
||||
@@ -0,0 +1,4 @@
|
||||
- generic [active] [ref=f15e1]:
|
||||
- progressbar "Loading" [ref=f15e4]
|
||||
- generic:
|
||||
- list
|
||||
@@ -0,0 +1 @@
|
||||
- generic [active] [ref=f18e1]: Not Found
|
||||
@@ -0,0 +1,151 @@
|
||||
- generic [active] [ref=e1]:
|
||||
- generic [ref=e4]:
|
||||
- link "Skip to main content" [ref=e5] [cursor=pointer]:
|
||||
- /url: "#pageContent"
|
||||
- banner [ref=e7]:
|
||||
- generic [ref=e8]:
|
||||
- link [ref=e10] [cursor=pointer]:
|
||||
- /url: /
|
||||
- img "Grafana" [ref=e11]
|
||||
- generic [ref=e14]:
|
||||
- button "Search or jump to..." [ref=e18] [cursor=pointer]
|
||||
- generic [ref=e19]: ctrl+k
|
||||
- generic [ref=e23]:
|
||||
- button "New" [ref=e24] [cursor=pointer]
|
||||
- button "Help" [ref=e30] [cursor=pointer]
|
||||
- button "News" [ref=e33] [cursor=pointer]
|
||||
- button "Profile" [ref=e36] [cursor=pointer]:
|
||||
- img "User avatar" [ref=e37]
|
||||
- generic [ref=e38]:
|
||||
- button "Open menu" [ref=e40] [cursor=pointer]
|
||||
- navigation "Breadcrumbs" [ref=e43]:
|
||||
- list [ref=e44]:
|
||||
- listitem [ref=e45]:
|
||||
- link "Home" [ref=e46] [cursor=pointer]:
|
||||
- /url: /
|
||||
- listitem [ref=e50]:
|
||||
- link "Explore" [ref=e51] [cursor=pointer]:
|
||||
- /url: /explore
|
||||
- listitem [ref=e55]:
|
||||
- generic "Prometheus" [ref=e56]
|
||||
- generic [ref=e57]:
|
||||
- generic [ref=e60]:
|
||||
- button "Copy shortened URL" [ref=e61] [cursor=pointer]
|
||||
- button "Open copy link options" [ref=e64] [cursor=pointer]
|
||||
- button "Toggle top search bar" [ref=e68] [cursor=pointer]
|
||||
- main [ref=e74]:
|
||||
- generic [ref=e76]:
|
||||
- heading "Explore" [level=1] [ref=e77]
|
||||
- generic [ref=e82]:
|
||||
- navigation "Explore toolbar" [ref=e84]:
|
||||
- navigation "Search links" [ref=e86]:
|
||||
- generic [ref=e87]:
|
||||
- button "Content outline" [expanded] [ref=e89] [cursor=pointer]:
|
||||
- generic [ref=e92]: Outline
|
||||
- generic [ref=e97] [cursor=pointer]:
|
||||
- img "Prometheus logo" [ref=e99]
|
||||
- textbox "Select a data source" [ref=e100]:
|
||||
- /placeholder: ""
|
||||
- generic [ref=e104]:
|
||||
- button "Split the pane" [ref=e106] [cursor=pointer]:
|
||||
- generic [ref=e109]: Split
|
||||
- button "Add" [ref=e111] [cursor=pointer]
|
||||
- generic [ref=e116]:
|
||||
- 'button "Time range selected: Last 45 minutes" [ref=e117] [cursor=pointer]'
|
||||
- button "Zoom out time range" [ref=e122] [cursor=pointer]
|
||||
- generic [ref=e126]:
|
||||
- button "Run query" [ref=e127] [cursor=pointer]
|
||||
- button "Auto refresh turned off. Choose refresh time interval" [ref=e131] [cursor=pointer]
|
||||
- generic [ref=e135]:
|
||||
- generic [ref=e139]:
|
||||
- button "Collapse outline" [expanded] [ref=e141] [cursor=pointer]:
|
||||
- img "arrow-from-right" [ref=e142]
|
||||
- button "Queries" [ref=e145] [cursor=pointer]:
|
||||
- img "arrow" [ref=e146]
|
||||
- button "Graph" [ref=e150] [cursor=pointer]:
|
||||
- img "graph-bar" [ref=e151]
|
||||
- generic [ref=e158]:
|
||||
- generic [ref=e160]:
|
||||
- generic "Query editor row" [ref=e163]:
|
||||
- generic [ref=e164]:
|
||||
- generic [ref=e166]:
|
||||
- generic [ref=e167]:
|
||||
- button "Collapse query row" [expanded] [ref=e168] [cursor=pointer]
|
||||
- generic [ref=e171]:
|
||||
- button "Query editor row title A" [ref=e172] [cursor=pointer]:
|
||||
- generic [ref=e173]: A
|
||||
- emphasis [ref=e174]: (Prometheus)
|
||||
- generic [ref=e175]:
|
||||
- button "Show data source help" [ref=e177] [cursor=pointer]
|
||||
- button "Duplicate query" [ref=e181] [cursor=pointer]
|
||||
- button "Hide response" [ref=e185] [cursor=pointer]
|
||||
- button "Remove query" [ref=e189] [cursor=pointer]
|
||||
- button "Drag and drop to reorder" [ref=e192]:
|
||||
- img "Drag and drop to reorder" [ref=e193]
|
||||
- generic [ref=e196]:
|
||||
- generic [ref=e197]:
|
||||
- button "Kick start your query" [ref=e198] [cursor=pointer]
|
||||
- generic [ref=e201]:
|
||||
- generic [ref=e202] [cursor=pointer]: Explain
|
||||
- generic [ref=e203]:
|
||||
- checkbox "Explain Toggle switch" [ref=e204]
|
||||
- generic "Toggle switch" [ref=e205] [cursor=pointer]
|
||||
- radiogroup [ref=e210]:
|
||||
- generic [ref=e211]:
|
||||
- radio "Builder" [ref=e212] [cursor=pointer]
|
||||
- generic [ref=e213] [cursor=pointer]: Builder
|
||||
- generic [ref=e214]:
|
||||
- radio "Code" [checked] [ref=e215] [cursor=pointer]
|
||||
- generic [ref=e216] [cursor=pointer]: Code
|
||||
- generic [ref=e218]:
|
||||
- generic [ref=e220]:
|
||||
- button "Metrics browser" [ref=e221] [cursor=pointer]
|
||||
- code [ref=e228]:
|
||||
- generic [ref=e229]:
|
||||
- generic [ref=e234]: vendor_cluster_size
|
||||
- textbox "Editor content;Press Alt+F1 for Accessibility Options." [ref=e239]: vendor_cluster_size
|
||||
- 'button "Options Legend: {{pod}} on {{node}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=e245] [cursor=pointer]':
|
||||
- generic [ref=e249]:
|
||||
- heading "Options" [level=6] [ref=e250]
|
||||
- generic [ref=e251]:
|
||||
- generic [ref=e252]: "Legend: {{pod}} on {{node}}"
|
||||
- generic [ref=e253]: "Format: Time series"
|
||||
- generic [ref=e254]: "Step: auto"
|
||||
- generic [ref=e255]: "Type: Range"
|
||||
- generic [ref=e256]: "Exemplars: false"
|
||||
- generic [ref=e257]:
|
||||
- button "Add query" [ref=e258] [cursor=pointer]
|
||||
- button "Query history" [ref=e262] [cursor=pointer]
|
||||
- button "Query inspector" [ref=e266] [cursor=pointer]
|
||||
- main [ref=e270]:
|
||||
- region [ref=e272]:
|
||||
- generic [ref=e273]:
|
||||
- heading "Graph" [level=2] [ref=e275]
|
||||
- radiogroup [ref=e278]:
|
||||
- generic [ref=e279]:
|
||||
- radio "Lines" [checked] [ref=e280] [cursor=pointer]
|
||||
- generic [ref=e281] [cursor=pointer]: Lines
|
||||
- generic [ref=e282]:
|
||||
- radio "Bars" [ref=e283] [cursor=pointer]
|
||||
- generic [ref=e284] [cursor=pointer]: Bars
|
||||
- generic [ref=e285]:
|
||||
- radio "Points" [ref=e286] [cursor=pointer]
|
||||
- generic [ref=e287] [cursor=pointer]: Points
|
||||
- generic [ref=e288]:
|
||||
- radio "Stacked lines" [ref=e289] [cursor=pointer]
|
||||
- generic [ref=e290] [cursor=pointer]: Stacked lines
|
||||
- generic [ref=e291]:
|
||||
- radio "Stacked bars" [ref=e292] [cursor=pointer]
|
||||
- generic [ref=e293] [cursor=pointer]: Stacked bars
|
||||
- list [ref=e302]:
|
||||
- listitem [ref=e303]:
|
||||
- button "keycloak-0 on kc-lab-2" [ref=e307] [cursor=pointer]
|
||||
- listitem [ref=e308]:
|
||||
- button "keycloak-0 on kc-lab-2" [ref=e312] [cursor=pointer]
|
||||
- listitem [ref=e313]:
|
||||
- button "keycloak-1 on kc-lab-1" [ref=e317] [cursor=pointer]
|
||||
- generic [ref=e322]:
|
||||
- alert
|
||||
- alert
|
||||
- complementary
|
||||
- complementary
|
||||
@@ -0,0 +1,145 @@
|
||||
- generic [active] [ref=f3e1]:
|
||||
- generic [ref=f3e4]:
|
||||
- link "Skip to main content" [ref=f3e5] [cursor=pointer]:
|
||||
- /url: "#pageContent"
|
||||
- banner [ref=f3e7]:
|
||||
- generic [ref=f3e8]:
|
||||
- link [ref=f3e10] [cursor=pointer]:
|
||||
- /url: /
|
||||
- img "Grafana" [ref=f3e11]
|
||||
- generic [ref=f3e14]:
|
||||
- button "Search or jump to..." [ref=f3e18] [cursor=pointer]
|
||||
- generic [ref=f3e19]: ctrl+k
|
||||
- generic [ref=f3e23]:
|
||||
- button "New" [ref=f3e24] [cursor=pointer]
|
||||
- button "Help" [ref=f3e30] [cursor=pointer]
|
||||
- button "News" [ref=f3e33] [cursor=pointer]
|
||||
- button "Profile" [ref=f3e36] [cursor=pointer]:
|
||||
- img "User avatar" [ref=f3e37]
|
||||
- generic [ref=f3e38]:
|
||||
- button "Open menu" [ref=f3e40] [cursor=pointer]
|
||||
- navigation "Breadcrumbs" [ref=f3e43]:
|
||||
- list [ref=f3e44]:
|
||||
- listitem [ref=f3e45]:
|
||||
- link "Home" [ref=f3e46] [cursor=pointer]:
|
||||
- /url: /
|
||||
- listitem [ref=f3e50]:
|
||||
- link "Explore" [ref=f3e51] [cursor=pointer]:
|
||||
- /url: /explore
|
||||
- listitem [ref=f3e55]:
|
||||
- generic "Prometheus" [ref=f3e56]
|
||||
- generic [ref=f3e57]:
|
||||
- generic [ref=f3e60]:
|
||||
- button "Copy shortened URL" [ref=f3e61] [cursor=pointer]
|
||||
- button "Open copy link options" [ref=f3e64] [cursor=pointer]
|
||||
- button "Toggle top search bar" [ref=f3e68] [cursor=pointer]
|
||||
- main [ref=f3e74]:
|
||||
- generic [ref=f3e76]:
|
||||
- heading "Explore" [level=1] [ref=f3e77]
|
||||
- generic [ref=f3e82]:
|
||||
- navigation "Explore toolbar" [ref=f3e84]:
|
||||
- navigation "Search links" [ref=f3e86]:
|
||||
- generic [ref=f3e87]:
|
||||
- button "Content outline" [expanded] [ref=f3e89] [cursor=pointer]:
|
||||
- generic [ref=f3e92]: Outline
|
||||
- generic [ref=f3e97] [cursor=pointer]:
|
||||
- img "Prometheus logo" [ref=f3e99]
|
||||
- textbox "Select a data source" [ref=f3e100]:
|
||||
- /placeholder: ""
|
||||
- generic [ref=f3e104]:
|
||||
- button "Split the pane" [ref=f3e106] [cursor=pointer]:
|
||||
- generic [ref=f3e109]: Split
|
||||
- button "Add" [ref=f3e111] [cursor=pointer]
|
||||
- generic [ref=f3e116]:
|
||||
- 'button "Time range selected: Last 20 minutes" [ref=f3e117] [cursor=pointer]'
|
||||
- button "Zoom out time range" [ref=f3e122] [cursor=pointer]
|
||||
- generic [ref=f3e126]:
|
||||
- button "Run query" [ref=f3e127] [cursor=pointer]
|
||||
- button "Auto refresh turned off. Choose refresh time interval" [ref=f3e131] [cursor=pointer]
|
||||
- generic [ref=f3e135]:
|
||||
- generic [ref=f3e139]:
|
||||
- button "Collapse outline" [expanded] [ref=f3e141] [cursor=pointer]:
|
||||
- img "arrow-from-right" [ref=f3e142]
|
||||
- button "Queries" [ref=f3e145] [cursor=pointer]:
|
||||
- img "arrow" [ref=f3e146]
|
||||
- button "Graph" [ref=f3e150] [cursor=pointer]:
|
||||
- img "graph-bar" [ref=f3e151]
|
||||
- generic [ref=f3e158]:
|
||||
- generic [ref=f3e160]:
|
||||
- generic "Query editor row" [ref=f3e163]:
|
||||
- generic [ref=f3e164]:
|
||||
- generic [ref=f3e166]:
|
||||
- generic [ref=f3e167]:
|
||||
- button "Collapse query row" [expanded] [ref=f3e168] [cursor=pointer]
|
||||
- generic [ref=f3e171]:
|
||||
- button "Query editor row title A" [ref=f3e172] [cursor=pointer]:
|
||||
- generic [ref=f3e173]: A
|
||||
- emphasis [ref=f3e174]: (Prometheus)
|
||||
- generic [ref=f3e175]:
|
||||
- button "Show data source help" [ref=f3e177] [cursor=pointer]
|
||||
- button "Duplicate query" [ref=f3e181] [cursor=pointer]
|
||||
- button "Hide response" [ref=f3e185] [cursor=pointer]
|
||||
- button "Remove query" [ref=f3e189] [cursor=pointer]
|
||||
- button "Drag and drop to reorder" [ref=f3e192]:
|
||||
- img "Drag and drop to reorder" [ref=f3e193]
|
||||
- generic [ref=f3e196]:
|
||||
- generic [ref=f3e197]:
|
||||
- button "Kick start your query" [ref=f3e198] [cursor=pointer]
|
||||
- generic [ref=f3e201]:
|
||||
- generic [ref=f3e202] [cursor=pointer]: Explain
|
||||
- generic [ref=f3e203]:
|
||||
- checkbox "Explain Toggle switch" [ref=f3e204]
|
||||
- generic "Toggle switch" [ref=f3e205] [cursor=pointer]
|
||||
- radiogroup [ref=f3e210]:
|
||||
- generic [ref=f3e211]:
|
||||
- radio "Builder" [ref=f3e212] [cursor=pointer]
|
||||
- generic [ref=f3e213] [cursor=pointer]: Builder
|
||||
- generic [ref=f3e214]:
|
||||
- radio "Code" [checked] [ref=f3e215] [cursor=pointer]
|
||||
- generic [ref=f3e216] [cursor=pointer]: Code
|
||||
- generic [ref=f3e218]:
|
||||
- generic [ref=f3e220]:
|
||||
- button "Loading metrics..." [disabled] [ref=f3e221] [cursor=pointer]
|
||||
- code [ref=f3e228]:
|
||||
- generic [ref=f3e229]:
|
||||
- generic [ref=f3e234]: "up{job=\"keycloak\"}"
|
||||
- textbox "Editor content;Press Alt+F1 for Accessibility Options." [ref=f3e239]: "up{job=\"keycloak\"}"
|
||||
- 'button "Options Legend: up — {{pod}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=f3e245] [cursor=pointer]':
|
||||
- generic [ref=f3e249]:
|
||||
- heading "Options" [level=6] [ref=f3e250]
|
||||
- generic [ref=f3e251]:
|
||||
- generic [ref=f3e252]: "Legend: up — {{pod}}"
|
||||
- generic [ref=f3e253]: "Format: Time series"
|
||||
- generic [ref=f3e254]: "Step: auto"
|
||||
- generic [ref=f3e255]: "Type: Range"
|
||||
- generic [ref=f3e256]: "Exemplars: false"
|
||||
- generic [ref=f3e257]:
|
||||
- button "Add query" [ref=f3e258] [cursor=pointer]
|
||||
- button "Query history" [ref=f3e262] [cursor=pointer]
|
||||
- button "Query inspector" [ref=f3e266] [cursor=pointer]
|
||||
- main [ref=f3e270]:
|
||||
- region [ref=f3e272]:
|
||||
- generic [ref=f3e273]:
|
||||
- heading "Graph" [level=2] [ref=f3e275]
|
||||
- radiogroup [ref=f3e278]:
|
||||
- generic [ref=f3e279]:
|
||||
- radio "Lines" [checked] [ref=f3e280] [cursor=pointer]
|
||||
- generic [ref=f3e281] [cursor=pointer]: Lines
|
||||
- generic [ref=f3e282]:
|
||||
- radio "Bars" [ref=f3e283] [cursor=pointer]
|
||||
- generic [ref=f3e284] [cursor=pointer]: Bars
|
||||
- generic [ref=f3e285]:
|
||||
- radio "Points" [ref=f3e286] [cursor=pointer]
|
||||
- generic [ref=f3e287] [cursor=pointer]: Points
|
||||
- generic [ref=f3e288]:
|
||||
- radio "Stacked lines" [ref=f3e289] [cursor=pointer]
|
||||
- generic [ref=f3e290] [cursor=pointer]: Stacked lines
|
||||
- generic [ref=f3e291]:
|
||||
- radio "Stacked bars" [ref=f3e292] [cursor=pointer]
|
||||
- generic [ref=f3e293] [cursor=pointer]: Stacked bars
|
||||
- generic [ref=f3e294]: Loading plugin panel...
|
||||
- generic [ref=f3e299]:
|
||||
- alert
|
||||
- alert
|
||||
- complementary
|
||||
- complementary
|
||||
@@ -0,0 +1,44 @@
|
||||
- main [ref=f6e7]:
|
||||
- generic [ref=f6e9]:
|
||||
- generic [ref=f6e11]:
|
||||
- generic [ref=f6e12]:
|
||||
- img "Grafana" [ref=f6e13]
|
||||
- heading "Welcome to Grafana" [level=1] [ref=f6e15]
|
||||
- generic [ref=f6e19]:
|
||||
- generic [ref=f6e20]:
|
||||
- generic [ref=f6e21]: Email or username
|
||||
- textbox "Email or username" [active] [ref=f6e28]:
|
||||
- /placeholder: email or username
|
||||
- generic [ref=f6e29]:
|
||||
- generic [ref=f6e30]: Password
|
||||
- generic [ref=f6e36]:
|
||||
- textbox "Password" [ref=f6e37]:
|
||||
- /placeholder: password
|
||||
- switch "Show password" [ref=f6e39] [cursor=pointer]
|
||||
- button "Log in" [ref=f6e42] [cursor=pointer]
|
||||
- link "Forgot your password?" [ref=f6e45] [cursor=pointer]:
|
||||
- /url: /user/password/send-reset-email
|
||||
- list [ref=f6e49]:
|
||||
- listitem [ref=f6e50]:
|
||||
- link "Documentation" [ref=f6e53] [cursor=pointer]:
|
||||
- /url: https://grafana.com/docs/grafana/latest/?utm_source=grafana_footer
|
||||
- text: "|"
|
||||
- listitem [ref=f6e54]:
|
||||
- link "Support" [ref=f6e57] [cursor=pointer]:
|
||||
- /url: https://grafana.com/products/enterprise/?utm_source=grafana_footer
|
||||
- text: "|"
|
||||
- listitem [ref=f6e58]:
|
||||
- link "Community" [ref=f6e61] [cursor=pointer]:
|
||||
- /url: https://community.grafana.com/?utm_source=grafana_footer
|
||||
- text: "|"
|
||||
- listitem [ref=f6e62]:
|
||||
- link "Open Source" [ref=f6e63] [cursor=pointer]:
|
||||
- /url: https://grafana.com/oss/grafana?utm_source=grafana_footer
|
||||
- text: "|"
|
||||
- listitem [ref=f6e64]:
|
||||
- link "Grafana v11.4.0 (b58701869e)" [ref=f6e65] [cursor=pointer]:
|
||||
- /url: https://github.com/grafana/grafana/blob/main/CHANGELOG.md
|
||||
- text: "|"
|
||||
- listitem [ref=f6e66]:
|
||||
- link "New version available!" [ref=f6e69] [cursor=pointer]:
|
||||
- /url: https://grafana.com/grafana/download?utm_source=grafana_footer
|
||||
@@ -0,0 +1,159 @@
|
||||
- generic [active] [ref=f9e1]:
|
||||
- generic [ref=f9e4]:
|
||||
- link "Skip to main content" [ref=f9e5] [cursor=pointer]:
|
||||
- /url: "#pageContent"
|
||||
- banner [ref=f9e7]:
|
||||
- generic [ref=f9e8]:
|
||||
- link [ref=f9e10] [cursor=pointer]:
|
||||
- /url: /
|
||||
- img "Grafana" [ref=f9e11]
|
||||
- generic [ref=f9e14]:
|
||||
- button "Search or jump to..." [ref=f9e18] [cursor=pointer]
|
||||
- generic [ref=f9e19]: ctrl+k
|
||||
- generic [ref=f9e23]:
|
||||
- button "New" [ref=f9e24] [cursor=pointer]
|
||||
- button "Help" [ref=f9e30] [cursor=pointer]
|
||||
- button "News" [ref=f9e33] [cursor=pointer]
|
||||
- button "Profile" [ref=f9e36] [cursor=pointer]:
|
||||
- img "User avatar" [ref=f9e37]
|
||||
- generic [ref=f9e38]:
|
||||
- button "Open menu" [ref=f9e40] [cursor=pointer]
|
||||
- navigation "Breadcrumbs" [ref=f9e43]:
|
||||
- list [ref=f9e44]:
|
||||
- listitem [ref=f9e45]:
|
||||
- link "Home" [ref=f9e46] [cursor=pointer]:
|
||||
- /url: /
|
||||
- listitem [ref=f9e50]:
|
||||
- link "Explore" [ref=f9e51] [cursor=pointer]:
|
||||
- /url: /explore
|
||||
- listitem [ref=f9e55]:
|
||||
- generic "Prometheus" [ref=f9e56]
|
||||
- generic [ref=f9e57]:
|
||||
- generic [ref=f9e60]:
|
||||
- button "Copy shortened URL" [ref=f9e61] [cursor=pointer]
|
||||
- button "Open copy link options" [ref=f9e64] [cursor=pointer]
|
||||
- button "Toggle top search bar" [ref=f9e68] [cursor=pointer]
|
||||
- main [ref=f9e74]:
|
||||
- generic [ref=f9e76]:
|
||||
- heading "Explore" [level=1] [ref=f9e77]
|
||||
- generic [ref=f9e82]:
|
||||
- navigation "Explore toolbar" [ref=f9e84]:
|
||||
- navigation "Search links" [ref=f9e86]:
|
||||
- generic [ref=f9e87]:
|
||||
- button "Content outline" [expanded] [ref=f9e89] [cursor=pointer]:
|
||||
- generic [ref=f9e92]: Outline
|
||||
- generic [ref=f9e97] [cursor=pointer]:
|
||||
- img "Prometheus logo" [ref=f9e99]
|
||||
- textbox "Select a data source" [ref=f9e100]:
|
||||
- /placeholder: ""
|
||||
- generic [ref=f9e104]:
|
||||
- button "Split the pane" [ref=f9e106] [cursor=pointer]:
|
||||
- generic [ref=f9e109]: Split
|
||||
- button "Add" [ref=f9e111] [cursor=pointer]
|
||||
- generic [ref=f9e116]:
|
||||
- 'button "Time range selected: Last 55 minutes" [ref=f9e117] [cursor=pointer]'
|
||||
- button "Zoom out time range" [ref=f9e122] [cursor=pointer]
|
||||
- generic [ref=f9e126]:
|
||||
- button "Run query" [ref=f9e127] [cursor=pointer]
|
||||
- button "Auto refresh turned off. Choose refresh time interval" [ref=f9e131] [cursor=pointer]
|
||||
- generic [ref=f9e135]:
|
||||
- generic [ref=f9e139]:
|
||||
- button "Collapse outline" [expanded] [ref=f9e141] [cursor=pointer]:
|
||||
- img "arrow-from-right" [ref=f9e142]
|
||||
- button "Queries" [ref=f9e145] [cursor=pointer]:
|
||||
- img "arrow" [ref=f9e146]
|
||||
- button "Graph" [ref=f9e150] [cursor=pointer]:
|
||||
- img "graph-bar" [ref=f9e151]
|
||||
- generic [ref=f9e158]:
|
||||
- generic [ref=f9e160]:
|
||||
- generic "Query editor row" [ref=f9e163]:
|
||||
- generic [ref=f9e164]:
|
||||
- generic [ref=f9e166]:
|
||||
- generic [ref=f9e167]:
|
||||
- button "Collapse query row" [expanded] [ref=f9e168] [cursor=pointer]
|
||||
- generic [ref=f9e171]:
|
||||
- button "Query editor row title A" [ref=f9e172] [cursor=pointer]:
|
||||
- generic [ref=f9e173]: A
|
||||
- emphasis [ref=f9e174]: (Prometheus)
|
||||
- generic [ref=f9e175]:
|
||||
- button "Show data source help" [ref=f9e177] [cursor=pointer]
|
||||
- button "Duplicate query" [ref=f9e181] [cursor=pointer]
|
||||
- button "Hide response" [ref=f9e185] [cursor=pointer]
|
||||
- button "Remove query" [ref=f9e189] [cursor=pointer]
|
||||
- button "Drag and drop to reorder" [ref=f9e192]:
|
||||
- img "Drag and drop to reorder" [ref=f9e193]
|
||||
- generic [ref=f9e196]:
|
||||
- generic [ref=f9e197]:
|
||||
- button "Kick start your query" [ref=f9e198] [cursor=pointer]
|
||||
- generic [ref=f9e201]:
|
||||
- generic [ref=f9e202] [cursor=pointer]: Explain
|
||||
- generic [ref=f9e203]:
|
||||
- checkbox "Explain Toggle switch" [ref=f9e204]
|
||||
- generic "Toggle switch" [ref=f9e205] [cursor=pointer]
|
||||
- radiogroup [ref=f9e210]:
|
||||
- generic [ref=f9e211]:
|
||||
- radio "Builder" [ref=f9e212] [cursor=pointer]
|
||||
- generic [ref=f9e213] [cursor=pointer]: Builder
|
||||
- generic [ref=f9e214]:
|
||||
- radio "Code" [checked] [ref=f9e215] [cursor=pointer]
|
||||
- generic [ref=f9e216] [cursor=pointer]: Code
|
||||
- generic [ref=f9e218]:
|
||||
- generic [ref=f9e220]:
|
||||
- button "Metrics browser" [ref=f9e221] [cursor=pointer]
|
||||
- code [ref=f9e228]:
|
||||
- generic [ref=f9e229]:
|
||||
- generic [ref=f9e234]: "up{job=~\"keycloak|node-exporter\"}"
|
||||
- textbox "Editor content;Press Alt+F1 for Accessibility Options." [ref=f9e239]: "up{job=~\"keycloak|node-exporter\"}"
|
||||
- 'button "Options Legend: {{job}} — {{pod}}{{node}} Format: Time series Step: auto Type: Range Exemplars: false" [ref=f9e245] [cursor=pointer]':
|
||||
- generic [ref=f9e249]:
|
||||
- heading "Options" [level=6] [ref=f9e250]
|
||||
- generic [ref=f9e251]:
|
||||
- generic [ref=f9e252]: "Legend: {{job}} — {{pod}}{{node}}"
|
||||
- generic [ref=f9e253]: "Format: Time series"
|
||||
- generic [ref=f9e254]: "Step: auto"
|
||||
- generic [ref=f9e255]: "Type: Range"
|
||||
- generic [ref=f9e256]: "Exemplars: false"
|
||||
- generic [ref=f9e257]:
|
||||
- button "Add query" [ref=f9e258] [cursor=pointer]
|
||||
- button "Query history" [ref=f9e262] [cursor=pointer]
|
||||
- button "Query inspector" [ref=f9e266] [cursor=pointer]
|
||||
- main [ref=f9e270]:
|
||||
- region [ref=f9e272]:
|
||||
- generic [ref=f9e273]:
|
||||
- heading "Graph" [level=2] [ref=f9e275]
|
||||
- radiogroup [ref=f9e278]:
|
||||
- generic [ref=f9e279]:
|
||||
- radio "Lines" [checked] [ref=f9e280] [cursor=pointer]
|
||||
- generic [ref=f9e281] [cursor=pointer]: Lines
|
||||
- generic [ref=f9e282]:
|
||||
- radio "Bars" [ref=f9e283] [cursor=pointer]
|
||||
- generic [ref=f9e284] [cursor=pointer]: Bars
|
||||
- generic [ref=f9e285]:
|
||||
- radio "Points" [ref=f9e286] [cursor=pointer]
|
||||
- generic [ref=f9e287] [cursor=pointer]: Points
|
||||
- generic [ref=f9e288]:
|
||||
- radio "Stacked lines" [ref=f9e289] [cursor=pointer]
|
||||
- generic [ref=f9e290] [cursor=pointer]: Stacked lines
|
||||
- generic [ref=f9e291]:
|
||||
- radio "Stacked bars" [ref=f9e292] [cursor=pointer]
|
||||
- generic [ref=f9e293] [cursor=pointer]: Stacked bars
|
||||
- list [ref=f9e311]:
|
||||
- listitem [ref=f9e312]:
|
||||
- button "keycloak — keycloak-1kc-lab-1" [ref=f9e316] [cursor=pointer]
|
||||
- listitem [ref=f9e317]:
|
||||
- button "keycloak — keycloak-1kc-lab-1" [ref=f9e321] [cursor=pointer]
|
||||
- listitem [ref=f9e322]:
|
||||
- button "keycloak — keycloak-0kc-lab-2" [ref=f9e326] [cursor=pointer]
|
||||
- listitem [ref=f9e327]:
|
||||
- button "keycloak — keycloak-0kc-lab-2" [ref=f9e331] [cursor=pointer]
|
||||
- listitem [ref=f9e332]:
|
||||
- button "keycloak — keycloak-0kc-lab-2" [ref=f9e336] [cursor=pointer]
|
||||
- listitem [ref=f9e337]:
|
||||
- button "node-exporter — kc-lab-1" [ref=f9e341] [cursor=pointer]
|
||||
- listitem [ref=f9e342]:
|
||||
- button "node-exporter — kc-lab-2" [ref=f9e346] [cursor=pointer]
|
||||
- generic [ref=f9e351]:
|
||||
- alert
|
||||
- alert
|
||||
- complementary
|
||||
- complementary
|
||||
@@ -0,0 +1,46 @@
|
||||
# Experiment A-1 — cut the JGroups transport (TCP 7800) while leaving discovery alone.
|
||||
#
|
||||
# The point is to separate two things that are easy to conflate:
|
||||
#
|
||||
# discovery how the nodes FIND each other -> PostgreSQL JGROUPS_PING table
|
||||
# transport how they actually TALK -> TCP 7800
|
||||
#
|
||||
# Blocking only the transport produces a state that cannot happen on a single
|
||||
# node: both members stay registered in the database, so each believes the other
|
||||
# exists, yet no message gets through.
|
||||
#
|
||||
# kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
|
||||
# kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
|
||||
#
|
||||
# NetworkPolicy is an ALLOWLIST, not a firewall with deny rules. There is no way
|
||||
# to write "deny 7800". The moment a pod is selected by a policy carrying
|
||||
# policyTypes: [Ingress], every inbound port is denied unless a rule permits it.
|
||||
# So 7800 is blocked by *omission*: 8080 and 9000 are listed, 7800 is not.
|
||||
#
|
||||
# That makes the two allow rules load-bearing — get them wrong and the experiment
|
||||
# measures a dead Keycloak instead of a partitioned cluster:
|
||||
#
|
||||
# 8080 the HTTP endpoint. Traefik, the other pod's REST calls, and the probe
|
||||
# traffic all arrive here.
|
||||
# 9000 the management port: /health/started, /health/ready, /health/live and
|
||||
# /metrics. Losing it means the kubelet fails the readiness probe and
|
||||
# kills the pod — the cluster would break for the wrong reason.
|
||||
#
|
||||
# Both rules deliberately omit `from:`, which allows those ports from any source.
|
||||
# Narrowing the source is not the subject here; the 2-hop experiment already
|
||||
# established how to do that by label when it matters.
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: a1-block-jgroups-transport
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app: keycloak
|
||||
policyTypes: [Ingress]
|
||||
ingress:
|
||||
- ports:
|
||||
- { port: 8080, protocol: TCP } # HTTP — must stay open
|
||||
- { port: 9000, protocol: TCP } # health + metrics — must stay open
|
||||
# 7800 is absent on purpose. That is the whole experiment.
|
||||
@@ -0,0 +1,277 @@
|
||||
# Keycloak multi-node cluster with PostgreSQL.
|
||||
#
|
||||
# Goal of this manifest: two Keycloak pods on two different nodes must discover
|
||||
# each other and form one Infinispan cluster. Keycloak 26 discovers peers through
|
||||
# the database (jdbc-ping) rather than multicast, writing to a JGROUPS_PING table,
|
||||
# but the cluster traffic itself runs over TCP 7800 between the pods. Those are
|
||||
# two separate mechanisms, which is why "registered in the DB but not clustered"
|
||||
# is a real failure mode — and one that a single node cannot reproduce.
|
||||
#
|
||||
# kubectl apply -f deploy/lab/k8s/keycloak-cluster.yaml
|
||||
# kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=600s
|
||||
#
|
||||
# Secrets are plain here. Proper secret handling is roadmap item 11; keeping it
|
||||
# visible for now is deliberate so the gap is obvious rather than forgotten.
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: keycloak-lab
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: keycloak-lab-secrets
|
||||
namespace: keycloak-lab
|
||||
type: Opaque
|
||||
stringData:
|
||||
POSTGRES_PASSWORD: lab-postgres-change-me
|
||||
KC_BOOTSTRAP_ADMIN_PASSWORD: lab-admin-change-me
|
||||
---
|
||||
# PostgreSQL. local-path binds the volume to whichever node the pod lands on, so
|
||||
# the database is effectively pinned to one node. That is not a flaw here: it is
|
||||
# what makes "the database node dies" a meaningful experiment later.
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: postgres-data
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: local-path
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: postgres
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: Recreate # RWO volume cannot be mounted by two pods at once
|
||||
selector:
|
||||
matchLabels:
|
||||
app: postgres
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: postgres
|
||||
spec:
|
||||
containers:
|
||||
- name: postgres
|
||||
image: postgres:16-alpine
|
||||
ports:
|
||||
- containerPort: 5432
|
||||
name: postgres
|
||||
env:
|
||||
- name: POSTGRES_DB
|
||||
value: keycloak
|
||||
- name: POSTGRES_USER
|
||||
value: keycloak
|
||||
- name: POSTGRES_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: keycloak-lab-secrets
|
||||
key: POSTGRES_PASSWORD
|
||||
# The image refuses to initialise into a non-empty mount, and
|
||||
# local-path volumes are clean, but this keeps the data one level
|
||||
# down so a lost+found or similar never blocks initdb.
|
||||
- name: PGDATA
|
||||
value: /var/lib/postgresql/data/pgdata
|
||||
volumeMounts:
|
||||
- name: data
|
||||
mountPath: /var/lib/postgresql/data
|
||||
readinessProbe:
|
||||
exec:
|
||||
command: ["sh", "-c", "pg_isready -U keycloak -d keycloak"]
|
||||
initialDelaySeconds: 10
|
||||
periodSeconds: 5
|
||||
resources:
|
||||
requests:
|
||||
memory: 192Mi
|
||||
cpu: 50m
|
||||
limits:
|
||||
memory: 512Mi
|
||||
volumes:
|
||||
- name: data
|
||||
persistentVolumeClaim:
|
||||
claimName: postgres-data
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: postgres
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
selector:
|
||||
app: postgres
|
||||
ports:
|
||||
- port: 5432
|
||||
targetPort: postgres
|
||||
---
|
||||
# Keycloak. A StatefulSet rather than a Deployment so each pod keeps a stable
|
||||
# name (keycloak-0, keycloak-1); cluster membership is far easier to read in
|
||||
# logs and in the JGROUPS_PING table when the identities do not churn.
|
||||
apiVersion: apps/v1
|
||||
kind: StatefulSet
|
||||
metadata:
|
||||
name: keycloak
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
serviceName: keycloak-headless
|
||||
replicas: 2
|
||||
podManagementPolicy: Parallel # both pods start together, so they race to
|
||||
# register — which is the interesting case
|
||||
selector:
|
||||
matchLabels:
|
||||
app: keycloak
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: keycloak
|
||||
spec:
|
||||
# One pod per node. Two pods on one node would share a kernel and make the
|
||||
# 7800 blocking experiment meaningless.
|
||||
topologySpreadConstraints:
|
||||
- maxSkew: 1
|
||||
topologyKey: kubernetes.io/hostname
|
||||
whenUnsatisfiable: ScheduleAnyway
|
||||
labelSelector:
|
||||
matchLabels:
|
||||
app: keycloak
|
||||
containers:
|
||||
- name: keycloak
|
||||
image: quay.io/keycloak/keycloak:26.7.0
|
||||
# "start", not "start-dev". Dev mode forces cache=local and there is
|
||||
# no cluster to form at all.
|
||||
args: ["start"]
|
||||
ports:
|
||||
- containerPort: 8080
|
||||
name: http
|
||||
- containerPort: 9000
|
||||
name: management
|
||||
- containerPort: 7800
|
||||
name: jgroups
|
||||
env:
|
||||
- name: KC_DB
|
||||
value: postgres
|
||||
- name: KC_DB_URL
|
||||
value: jdbc:postgresql://postgres:5432/keycloak
|
||||
- name: KC_DB_USERNAME
|
||||
value: keycloak
|
||||
- name: KC_DB_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: keycloak-lab-secrets
|
||||
key: POSTGRES_PASSWORD
|
||||
|
||||
# Settings confirmed by the two-hop header measurement.
|
||||
# KC_HOSTNAME carries the full external URL, which pins scheme and
|
||||
# host for issuer and redirect URLs regardless of headers.
|
||||
# KC_PROXY_HEADERS is the separate opt-in that lets the forwarded
|
||||
# client address through — the same kind of switch as Spring's
|
||||
# forward-headers-strategy. See docs/two-hop-proxy-header-contract.md.
|
||||
- name: KC_HOSTNAME
|
||||
value: https://auth.hyeonworks.com
|
||||
- name: KC_HOSTNAME_STRICT
|
||||
value: "true"
|
||||
- name: KC_PROXY_HEADERS
|
||||
value: xforwarded
|
||||
- name: KC_HTTP_ENABLED
|
||||
value: "true"
|
||||
|
||||
- name: KC_HEALTH_ENABLED
|
||||
value: "true"
|
||||
- name: KC_METRICS_ENABLED
|
||||
value: "true"
|
||||
|
||||
# Without an explicit cap the JVM sizes its heap from the container
|
||||
# limit and this lab has roughly 3.8GB of guest headroom in total.
|
||||
- name: JAVA_OPTS_KC_HEAP
|
||||
value: "-Xms256m -Xmx512m"
|
||||
|
||||
- name: KC_BOOTSTRAP_ADMIN_USERNAME
|
||||
value: admin
|
||||
- name: KC_BOOTSTRAP_ADMIN_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: keycloak-lab-secrets
|
||||
key: KC_BOOTSTRAP_ADMIN_PASSWORD
|
||||
|
||||
# Keycloak serves health and metrics on the management port (9000),
|
||||
# not on 8080, since version 25.
|
||||
startupProbe:
|
||||
httpGet:
|
||||
path: /health/started
|
||||
port: management
|
||||
periodSeconds: 10
|
||||
failureThreshold: 60 # first boot runs an implicit build
|
||||
readinessProbe:
|
||||
httpGet:
|
||||
path: /health/ready
|
||||
port: management
|
||||
periodSeconds: 10
|
||||
livenessProbe:
|
||||
httpGet:
|
||||
path: /health/live
|
||||
port: management
|
||||
periodSeconds: 30
|
||||
resources:
|
||||
requests:
|
||||
memory: 640Mi
|
||||
cpu: 100m
|
||||
limits:
|
||||
memory: 900Mi
|
||||
---
|
||||
# Headless service. Not required for jdbc-ping discovery, which goes through the
|
||||
# database, but it gives each pod a stable DNS name for direct inspection.
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: keycloak-headless
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
clusterIP: None
|
||||
selector:
|
||||
app: keycloak
|
||||
ports:
|
||||
- port: 8080
|
||||
targetPort: http
|
||||
name: http
|
||||
- port: 9000
|
||||
targetPort: management
|
||||
name: management
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: keycloak
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
selector:
|
||||
app: keycloak
|
||||
ports:
|
||||
- port: 8080
|
||||
targetPort: http
|
||||
name: http
|
||||
---
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: Ingress
|
||||
metadata:
|
||||
name: keycloak
|
||||
namespace: keycloak-lab
|
||||
spec:
|
||||
ingressClassName: traefik
|
||||
rules:
|
||||
- host: auth.hyeonworks.com
|
||||
http:
|
||||
paths:
|
||||
- path: /
|
||||
pathType: Prefix
|
||||
backend:
|
||||
service:
|
||||
name: keycloak
|
||||
port:
|
||||
number: 8080
|
||||
@@ -0,0 +1,373 @@
|
||||
# Prometheus + node-exporter + Grafana.
|
||||
#
|
||||
# Purpose: during a fault-injection experiment, know *which signal moved first*.
|
||||
# Without a metrics store the only record is whatever scrolled past in a terminal,
|
||||
# and "the cluster recovered in about a minute" is not a measurement.
|
||||
#
|
||||
# kubectl apply -f deploy/lab/k8s/observability.yaml
|
||||
# kubectl -n observability rollout status deployment/prometheus --timeout=300s
|
||||
#
|
||||
# Placement decision — Prometheus and Grafana are pinned to the control-plane
|
||||
# node (kc-lab-1). An observability stack must not share a failure domain with
|
||||
# the thing it observes. With only two nodes that cannot be fully avoided, so the
|
||||
# rule here is: the node that gets killed in experiments is the *agent*
|
||||
# (kc-lab-2, holding keycloak-0 and postgres), and everything needed to watch
|
||||
# that happen lives on the server node.
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: observability
|
||||
---
|
||||
# Prometheus discovers scrape targets by querying the Kubernetes API, so it
|
||||
# needs read access to nodes, services, endpoints and pods. Without this the
|
||||
# kubernetes_sd_configs below silently return no targets.
|
||||
apiVersion: v1
|
||||
kind: ServiceAccount
|
||||
metadata:
|
||||
name: prometheus
|
||||
namespace: observability
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
name: prometheus
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
# nodes/proxy is required in addition to nodes/metrics: the kubelet job
|
||||
# reaches each node through the API server's proxy subresource
|
||||
# (/api/v1/nodes/<name>/proxy/metrics). Without it every kubelet target
|
||||
# fails with 403 Forbidden while the other jobs stay green — a partial
|
||||
# failure that is easy to miss unless the target list is checked.
|
||||
resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
|
||||
verbs: [get, list, watch]
|
||||
- nonResourceURLs: ["/metrics"]
|
||||
verbs: [get]
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRoleBinding
|
||||
metadata:
|
||||
name: prometheus
|
||||
roleRef:
|
||||
apiGroup: rbac.authorization.k8s.io
|
||||
kind: ClusterRole
|
||||
name: prometheus
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: prometheus
|
||||
namespace: observability
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: prometheus-config
|
||||
namespace: observability
|
||||
data:
|
||||
prometheus.yml: |
|
||||
global:
|
||||
# 15s is short for production but right here: a node loss should show up
|
||||
# within a couple of samples, not a minute later.
|
||||
scrape_interval: 15s
|
||||
evaluation_interval: 15s
|
||||
|
||||
scrape_configs:
|
||||
# Prometheus scraping itself. Useful as a control: if this target is down,
|
||||
# the problem is Prometheus, not the thing being measured.
|
||||
- job_name: prometheus
|
||||
static_configs:
|
||||
- targets: ['localhost:9090']
|
||||
|
||||
# Keycloak. Metrics live on the management port 9000, not 8080 — the same
|
||||
# split that the health probes use. KC_METRICS_ENABLED=true is already set
|
||||
# on the StatefulSet.
|
||||
#
|
||||
# Discovery is by endpoints rather than a static list because pod IPs
|
||||
# change on every restart; that was observed directly when the lab was
|
||||
# power-cycled and every pod came back with a new address.
|
||||
- job_name: keycloak
|
||||
kubernetes_sd_configs:
|
||||
- role: endpoints
|
||||
namespaces:
|
||||
names: [keycloak-lab]
|
||||
relabel_configs:
|
||||
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
|
||||
action: keep
|
||||
regex: keycloak-headless;management
|
||||
- source_labels: [__meta_kubernetes_pod_name]
|
||||
target_label: pod
|
||||
- source_labels: [__meta_kubernetes_pod_node_name]
|
||||
target_label: node
|
||||
|
||||
# node-exporter, one per node via DaemonSet. This is what answers
|
||||
# "did the machine die or did the process die".
|
||||
- job_name: node-exporter
|
||||
kubernetes_sd_configs:
|
||||
- role: endpoints
|
||||
namespaces:
|
||||
names: [observability]
|
||||
relabel_configs:
|
||||
- source_labels: [__meta_kubernetes_service_name]
|
||||
action: keep
|
||||
regex: node-exporter
|
||||
- source_labels: [__meta_kubernetes_pod_node_name]
|
||||
target_label: node
|
||||
|
||||
# The kubelet's own metrics, reached through the API server proxy so no
|
||||
# extra port needs opening.
|
||||
- job_name: kubelet
|
||||
scheme: https
|
||||
tls_config:
|
||||
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
|
||||
insecure_skip_verify: true
|
||||
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
|
||||
kubernetes_sd_configs:
|
||||
- role: node
|
||||
relabel_configs:
|
||||
- action: labelmap
|
||||
regex: __meta_kubernetes_node_label_(.+)
|
||||
- target_label: __address__
|
||||
replacement: kubernetes.default.svc:443
|
||||
- source_labels: [__meta_kubernetes_node_name]
|
||||
regex: (.+)
|
||||
target_label: __metrics_path__
|
||||
replacement: /api/v1/nodes/${1}/proxy/metrics
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: prometheus-data
|
||||
namespace: observability
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: local-path
|
||||
resources:
|
||||
requests:
|
||||
storage: 5Gi
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: prometheus
|
||||
namespace: observability
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy:
|
||||
type: Recreate # RWO volume; two pods cannot mount it at once
|
||||
selector:
|
||||
matchLabels:
|
||||
app: prometheus
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: prometheus
|
||||
spec:
|
||||
serviceAccountName: prometheus
|
||||
# See the placement note at the top of this file.
|
||||
nodeSelector:
|
||||
node-role.kubernetes.io/control-plane: "true"
|
||||
securityContext:
|
||||
fsGroup: 65534 # the image runs as nobody and must own the volume
|
||||
containers:
|
||||
- name: prometheus
|
||||
image: prom/prometheus:v3.1.0
|
||||
args:
|
||||
- --config.file=/etc/prometheus/prometheus.yml
|
||||
- --storage.tsdb.path=/prometheus
|
||||
# 7 days is far more than an experiment needs and keeps the volume
|
||||
# small enough that it never becomes the reason a node fills up.
|
||||
- --storage.tsdb.retention.time=7d
|
||||
- --web.enable-lifecycle
|
||||
ports:
|
||||
- containerPort: 9090
|
||||
name: http
|
||||
volumeMounts:
|
||||
- name: config
|
||||
mountPath: /etc/prometheus
|
||||
- name: data
|
||||
mountPath: /prometheus
|
||||
readinessProbe:
|
||||
httpGet: { path: /-/ready, port: http }
|
||||
initialDelaySeconds: 10
|
||||
livenessProbe:
|
||||
httpGet: { path: /-/healthy, port: http }
|
||||
initialDelaySeconds: 30
|
||||
resources:
|
||||
requests: { memory: 256Mi, cpu: 50m }
|
||||
limits: { memory: 640Mi }
|
||||
volumes:
|
||||
- name: config
|
||||
configMap:
|
||||
name: prometheus-config
|
||||
- name: data
|
||||
persistentVolumeClaim:
|
||||
claimName: prometheus-data
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: prometheus
|
||||
namespace: observability
|
||||
spec:
|
||||
selector:
|
||||
app: prometheus
|
||||
ports:
|
||||
- port: 9090
|
||||
targetPort: http
|
||||
---
|
||||
# node-exporter. A DaemonSet so every node reports, including one that is about
|
||||
# to be killed — the last samples before it goes silent are the interesting part.
|
||||
apiVersion: apps/v1
|
||||
kind: DaemonSet
|
||||
metadata:
|
||||
name: node-exporter
|
||||
namespace: observability
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: node-exporter
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: node-exporter
|
||||
spec:
|
||||
# Host namespaces: the point is to measure the machine, not the container.
|
||||
hostNetwork: true
|
||||
hostPID: true
|
||||
tolerations:
|
||||
- operator: Exists # must also run on tainted nodes
|
||||
containers:
|
||||
- name: node-exporter
|
||||
image: prom/node-exporter:v1.8.2
|
||||
args:
|
||||
- --path.procfs=/host/proc
|
||||
- --path.sysfs=/host/sys
|
||||
- --path.rootfs=/host/root
|
||||
- --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
|
||||
ports:
|
||||
- containerPort: 9100
|
||||
name: metrics
|
||||
hostPort: 9100
|
||||
volumeMounts:
|
||||
- { name: proc, mountPath: /host/proc, readOnly: true }
|
||||
- { name: sys, mountPath: /host/sys, readOnly: true }
|
||||
- { name: rootfs, mountPath: /host/root, readOnly: true, mountPropagation: HostToContainer }
|
||||
resources:
|
||||
requests: { memory: 32Mi, cpu: 20m }
|
||||
limits: { memory: 96Mi }
|
||||
volumes:
|
||||
- { name: proc, hostPath: { path: /proc } }
|
||||
- { name: sys, hostPath: { path: /sys } }
|
||||
- { name: rootfs, hostPath: { path: / } }
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: node-exporter
|
||||
namespace: observability
|
||||
spec:
|
||||
clusterIP: None # headless: Prometheus wants each pod, not a VIP
|
||||
selector:
|
||||
app: node-exporter
|
||||
ports:
|
||||
- port: 9100
|
||||
targetPort: metrics
|
||||
name: metrics
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: grafana
|
||||
namespace: observability
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: grafana
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: grafana
|
||||
spec:
|
||||
nodeSelector:
|
||||
node-role.kubernetes.io/control-plane: "true"
|
||||
containers:
|
||||
- name: grafana
|
||||
image: grafana/grafana:11.4.0
|
||||
ports:
|
||||
- containerPort: 3000
|
||||
name: http
|
||||
env:
|
||||
- name: GF_SECURITY_ADMIN_USER
|
||||
value: admin
|
||||
- name: GF_SECURITY_ADMIN_PASSWORD
|
||||
value: lab-grafana-change-me
|
||||
# Grafana builds absolute URLs for redirects and asset paths. Behind
|
||||
# the nginx -> Traefik chain it must be told the external address,
|
||||
# for exactly the reason Keycloak needs KC_HOSTNAME. Without it,
|
||||
# login redirects come back as http://<pod-ip>:3000.
|
||||
- name: GF_SERVER_ROOT_URL
|
||||
value: https://app2.hyeonworks.com
|
||||
volumeMounts:
|
||||
- name: datasources
|
||||
mountPath: /etc/grafana/provisioning/datasources
|
||||
readinessProbe:
|
||||
httpGet: { path: /api/health, port: http }
|
||||
initialDelaySeconds: 15
|
||||
resources:
|
||||
requests: { memory: 128Mi, cpu: 50m }
|
||||
limits: { memory: 320Mi }
|
||||
volumes:
|
||||
- name: datasources
|
||||
configMap:
|
||||
name: grafana-datasources
|
||||
---
|
||||
# Provisioning the datasource as a file means Grafana comes up already wired to
|
||||
# Prometheus. Clicking through the UI would leave the configuration only in
|
||||
# Grafana's own database, which is emptyDir here and disappears on restart.
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: grafana-datasources
|
||||
namespace: observability
|
||||
data:
|
||||
prometheus.yaml: |
|
||||
apiVersion: 1
|
||||
datasources:
|
||||
- name: Prometheus
|
||||
type: prometheus
|
||||
access: proxy
|
||||
url: http://prometheus.observability.svc:9090
|
||||
isDefault: true
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: grafana
|
||||
namespace: observability
|
||||
spec:
|
||||
selector:
|
||||
app: grafana
|
||||
ports:
|
||||
- port: 3000
|
||||
targetPort: http
|
||||
---
|
||||
# Grafana is published on app2.hyeonworks.com because that name is already in
|
||||
# the wildcard-free certificate (auth / app1 / app2) and is otherwise unused.
|
||||
# It moves when app2 is needed for the SSO experiment.
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: Ingress
|
||||
metadata:
|
||||
name: grafana
|
||||
namespace: observability
|
||||
spec:
|
||||
ingressClassName: traefik
|
||||
rules:
|
||||
- host: app2.hyeonworks.com
|
||||
http:
|
||||
paths:
|
||||
- path: /
|
||||
pathType: Prefix
|
||||
backend:
|
||||
service:
|
||||
name: grafana
|
||||
port:
|
||||
number: 3000
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
#!/usr/bin/env bash
|
||||
# Experiment 0c — where does a session entry actually live?
|
||||
#
|
||||
# Experiment 0b showed keycloak-1's session cache never moved when keycloak-0
|
||||
# handled a login. That leaves two explanations:
|
||||
#
|
||||
# (a) a DISTRIBUTED cache with owners=1 — entries are spread across nodes by
|
||||
# consistent hashing, and this one happened to land on keycloak-0;
|
||||
# (b) a LOCAL cache — each node only ever caches what it handled itself.
|
||||
#
|
||||
# They are distinguished by driving logins at the OTHER node. Under (a) the
|
||||
# entries would keep landing on both nodes regardless of who was asked. Under
|
||||
# (b) the count rises only on the node that received the request.
|
||||
set -uo pipefail
|
||||
|
||||
NS="${NS:-keycloak-lab}"
|
||||
N="${N:-5}"
|
||||
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
|
||||
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
|
||||
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
|
||||
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
|
||||
|
||||
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
|
||||
echo " keycloak-0 = $K0_IP ($(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.spec.nodeName}'))"
|
||||
echo " keycloak-1 = $K1_IP ($(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.spec.nodeName}'))"
|
||||
echo
|
||||
|
||||
kubectl -n "$NS" run kc-own --rm -i --restart=Never \
|
||||
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
|
||||
O=/tmp/o; : > \$O
|
||||
ent() {
|
||||
curl -s --retry 3 --max-time 20 http://\$1:9000/metrics \
|
||||
| grep -E '^vendor_statistics_approximate_entries_unique.cache=.sessions' \
|
||||
| awk '{print \$NF}'
|
||||
}
|
||||
login() { i=0; while [ \$i -lt $N ]; do
|
||||
curl -s -o /dev/null -X POST http://\$1:8080/realms/master/protocol/openid-connect/token \
|
||||
-d grant_type=password -d client_id=admin-cli \
|
||||
-d username=admin -d 'password=$ADMIN_PW'
|
||||
i=\$((i+1)); done; sleep 5; }
|
||||
{
|
||||
printf '%-32s %12s %12s\n' '단계' 'k0 entries' 'k1 entries'
|
||||
printf '%-32s %12s %12s\n' '시작' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
|
||||
login $K1_IP
|
||||
printf '%-32s %12s %12s\n' 'keycloak-1 에 로그인 ${N}회' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
|
||||
login $K0_IP
|
||||
printf '%-32s %12s %12s\n' 'keycloak-0 에 로그인 ${N}회' \"\$(ent $K0_IP)\" \"\$(ent $K1_IP)\"
|
||||
} >> \$O
|
||||
cat \$O
|
||||
" 2>&1 | grep -v '^pod .* deleted$'
|
||||
|
||||
echo
|
||||
echo "=== 대조: PostgreSQL 에는 몇 건인가 ==="
|
||||
kubectl -n "$NS" exec deploy/postgres -- psql -U keycloak -d keycloak -tAc \
|
||||
"select count(*) from offline_user_session where offline_flag='0'" 2>/dev/null | sed 's/^/ online 세션 /'
|
||||
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env bash
|
||||
# Experiment 0b — does the Infinispan cache itself replicate, or do both nodes
|
||||
# merely agree because they read the same database?
|
||||
#
|
||||
# Experiment 0 proved the two nodes give the same answers. That alone does NOT
|
||||
# prove Infinispan replicated anything: with persistent-user-sessions (the
|
||||
# Keycloak 26 default) the session is written to PostgreSQL, so two nodes reading
|
||||
# one database would agree even with the cache disabled entirely.
|
||||
#
|
||||
# This script separates the two by measuring the cache counters on BOTH nodes
|
||||
# around a single login. If the write on keycloak-0 shows up as cache activity
|
||||
# on keycloak-1, the replication is real and not a database artifact.
|
||||
set -uo pipefail
|
||||
|
||||
NS="${NS:-keycloak-lab}"
|
||||
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
|
||||
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
|
||||
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
|
||||
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
|
||||
|
||||
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
|
||||
echo
|
||||
|
||||
# 파드 출력을 스트리밍으로 받으면 조각이 유실된다. 실제로 첫 시도에서
|
||||
# keycloak-1 의 스냅샷과 그 다음 마커가 통째로 사라져 델타가 0 으로 보였다.
|
||||
# 파드 안에서 파일로 모았다가 마지막에 한 번만 내보낸다.
|
||||
kubectl -n "$NS" run kc-delta --rm -i --restart=Never \
|
||||
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
|
||||
set -u
|
||||
K0='http://$K0_IP'; K1='http://$K1_IP'
|
||||
O=/tmp/o.txt; : > \$O
|
||||
snap() {
|
||||
curl -s --retry 3 --retry-connrefused --max-time 20 \$1:9000/metrics \
|
||||
| grep -E '^vendor_(statistics_(stores|hits|misses|approximate_entries_unique)|rpc_manager_replication_count)\{cache=\"(sessions|clientSessions)\"' \
|
||||
| sed 's/,cache_manager=\"keycloak\"//; s/,node=\"[^\"]*\"//' >> \$O
|
||||
}
|
||||
echo '###BEFORE_K0' >> \$O; snap \$K0
|
||||
echo '###BEFORE_K1' >> \$O; snap \$K1
|
||||
echo '###LOGIN' >> \$O
|
||||
curl -s -o /dev/null -w 'http_code=%{http_code}\n' -X POST \
|
||||
\"\$K0:8080/realms/master/protocol/openid-connect/token\" \
|
||||
-d grant_type=password -d client_id=admin-cli \
|
||||
-d username=admin -d 'password=$ADMIN_PW' >> \$O
|
||||
sleep 5
|
||||
echo '###AFTER_K0' >> \$O; snap \$K0
|
||||
echo '###AFTER_K1' >> \$O; snap \$K1
|
||||
echo '###END' >> \$O
|
||||
cat \$O
|
||||
" 2>&1 | grep -v '^pod .* deleted$' > /tmp/cache-delta.txt
|
||||
|
||||
python3 - /tmp/cache-delta.txt <<'PY'
|
||||
import re, sys
|
||||
raw = open(sys.argv[1]).read()
|
||||
blocks, cur = {}, None
|
||||
for line in raw.splitlines():
|
||||
if line.startswith('###'):
|
||||
cur = line[3:]; blocks[cur] = {}
|
||||
elif cur and '{' in line:
|
||||
m = re.match(r'(\S+?)\{cache="(\w+)"\}\s+(\S+)', line)
|
||||
if m:
|
||||
blocks[cur][(m.group(1), m.group(2))] = float(m.group(3))
|
||||
|
||||
print('=== 로그인은 keycloak-0 에만 보냈다 ===')
|
||||
code = [l for l in raw.splitlines() if l.startswith('http_code=')]
|
||||
print(' 로그인 응답: ' + (code[0] if code else '없음'))
|
||||
for n in ('BEFORE_K0','BEFORE_K1','AFTER_K0','AFTER_K1'):
|
||||
if not blocks.get(n):
|
||||
print(f' !! {n} 스냅샷이 비었다 — 델타를 신뢰할 수 없다')
|
||||
print()
|
||||
hdr = f" {'계수기':<42} {'캐시':<15} {'전':>8} {'후':>8} {'증가':>7}"
|
||||
for node in ('K0', 'K1'):
|
||||
who = 'keycloak-0 (로그인을 받은 노드)' if node == 'K0' else 'keycloak-1 (아무 요청도 받지 않은 노드)'
|
||||
print(f'=== {who} ===')
|
||||
print(hdr)
|
||||
b, a = blocks.get(f'BEFORE_{node}', {}), blocks.get(f'AFTER_{node}', {})
|
||||
for k in sorted(set(b) | set(a)):
|
||||
before, after = b.get(k[0:2], 0.0), a.get(k[0:2], 0.0)
|
||||
d = after - before
|
||||
mark = ' ←' if d else ''
|
||||
name = k[0].replace('vendor_statistics_', '').replace('vendor_rpc_manager_', 'rpc.')
|
||||
print(f" {name:<42} {k[1]:<15} {before:>8.0f} {after:>8.0f} {d:>+7.0f}{mark}")
|
||||
print()
|
||||
PY
|
||||
+108
@@ -0,0 +1,108 @@
|
||||
#!/usr/bin/env bash
|
||||
# Experiment 0d — capture the actual SQL that the OTHER node runs.
|
||||
#
|
||||
# Experiments 0b/0c showed that session entries never appear in keycloak-1's
|
||||
# memory, yet keycloak-1 can use a session keycloak-0 created. The conclusion
|
||||
# "keycloak-1 reads it from PostgreSQL" was an inference, not an observation.
|
||||
#
|
||||
# This script turns on statement logging in PostgreSQL for a few seconds, sends
|
||||
# ONE refresh request to keycloak-1 for a session born on keycloak-0, and greps
|
||||
# the database log for that session id. If the inference is right, the SQL is
|
||||
# there, issued from keycloak-1's pod IP.
|
||||
#
|
||||
# It also checks whether serving that request makes keycloak-1 cache the session
|
||||
# — which sharpens "each node caches what it handled" from "what it logged in"
|
||||
# to "what it touched".
|
||||
set -uo pipefail
|
||||
|
||||
NS="${NS:-keycloak-lab}"
|
||||
PSQL="kubectl -n $NS exec deploy/postgres -- psql -U keycloak -d keycloak -tAc"
|
||||
|
||||
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
|
||||
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
|
||||
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
|
||||
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
|
||||
|
||||
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
|
||||
echo " keycloak-0 = $K0_IP (세션을 만드는 노드)"
|
||||
echo " keycloak-1 = $K1_IP (읽기만 하는 노드)"
|
||||
echo
|
||||
|
||||
# %h 를 넣어야 어느 파드가 보낸 질의인지 로그에서 구분된다.
|
||||
echo "=== PostgreSQL 문장 로깅을 켠다 ==="
|
||||
$PSQL "alter system set log_statement='all'" >/dev/null 2>&1
|
||||
$PSQL "alter system set log_line_prefix='%m [%p] %h '" >/dev/null 2>&1
|
||||
$PSQL "select pg_reload_conf()" >/dev/null 2>&1
|
||||
echo " log_statement = $($PSQL 'show log_statement' 2>/dev/null)"
|
||||
echo " log_line_prefix = $($PSQL 'show log_line_prefix' 2>/dev/null)"
|
||||
echo
|
||||
|
||||
# 로그 커서를 잡아둔다. 이 줄 수 이후만 본다.
|
||||
LOG_BEFORE=$(kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null | wc -l)
|
||||
|
||||
RESULT=$(kubectl -n "$NS" run kc-readpath --rm -i --restart=Never \
|
||||
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
|
||||
O=/tmp/o; : > \$O
|
||||
TOKEN_EP='/realms/master/protocol/openid-connect/token'
|
||||
jget() { sed -n \"s/.*\\\"\$1\\\":\\\"\\([^\\\"]*\\)\\\".*/\\1/p\"; }
|
||||
ent() {
|
||||
curl -s --retry 3 --max-time 20 http://\$1:9000/metrics \
|
||||
| grep -E '^vendor_statistics_approximate_entries_unique.cache=.sessions' | awk '{print \$NF}'
|
||||
}
|
||||
# keycloak-0 에서 로그인한다
|
||||
L=\$(curl -s -X POST \"http://$K0_IP:8080\$TOKEN_EP\" -d grant_type=password \
|
||||
-d client_id=admin-cli -d username=admin -d 'password=$ADMIN_PW')
|
||||
SID=\$(echo \"\$L\" | jget access_token | cut -d. -f2 | sed 's/\$/==/' | base64 -d 2>/dev/null | jget sid)
|
||||
RT=\$(echo \"\$L\" | jget refresh_token)
|
||||
echo \"SID=\$SID\" >> \$O
|
||||
echo \"K1_ENTRIES_BEFORE=\$(ent $K1_IP)\" >> \$O
|
||||
sleep 2
|
||||
# 반대편 노드에 refresh 를 딱 한 번 보낸다
|
||||
# 인용을 한 겹 더 쌓으면 curl 이 URL 을 통째로 못 읽는다. 실제로 000 이 나왔다.
|
||||
CODE=\$(curl -s -o /dev/null -w '%{http_code}' -X POST \
|
||||
\"http://$K1_IP:8080\$TOKEN_EP\" \
|
||||
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\")
|
||||
echo \"REFRESH_ON_K1=\$CODE\" >> \$O
|
||||
sleep 3
|
||||
echo \"K1_ENTRIES_AFTER=\$(ent $K1_IP)\" >> \$O
|
||||
cat \$O
|
||||
" 2>&1 | grep -v '^pod .* deleted$')
|
||||
|
||||
echo "=== 요청 ==="
|
||||
echo "$RESULT" | sed 's/^/ /'
|
||||
SID=$(echo "$RESULT" | sed -n 's/^SID=//p')
|
||||
|
||||
echo
|
||||
echo "=== PostgreSQL 문장 로깅을 끈다 ==="
|
||||
$PSQL "alter system reset log_statement" >/dev/null 2>&1
|
||||
$PSQL "alter system reset log_line_prefix" >/dev/null 2>&1
|
||||
$PSQL "select pg_reload_conf()" >/dev/null 2>&1
|
||||
echo " log_statement = $($PSQL 'show log_statement' 2>/dev/null)"
|
||||
|
||||
echo
|
||||
echo "=== keycloak-1 이 실제로 보낸 SQL 문장 ==="
|
||||
echo " (파라미터가 \$1 로 묶여 있어, sid 는 바로 아래 DETAIL 줄에 있다)"
|
||||
echo
|
||||
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
|
||||
| tail -n +$((LOG_BEFORE + 1)) \
|
||||
| grep -F "$K1_IP" | grep -E "LOG: execute" \
|
||||
| sed 's/.*execute [^:]*: //' | sed 's/^/ /' | head -12
|
||||
echo
|
||||
echo "=== 그 sid 를 언급한 SQL — 누가 보냈는가 ==="
|
||||
echo " 찾는 sid: $SID"
|
||||
echo
|
||||
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
|
||||
| tail -n +$((LOG_BEFORE + 1)) \
|
||||
| grep -F "$SID" \
|
||||
| sed -e "s/$K0_IP/[keycloak-0]/g" -e "s/$K1_IP/[keycloak-1]/g" \
|
||||
| cut -c1-220 \
|
||||
| head -20
|
||||
|
||||
echo
|
||||
echo "=== 요약: 파드별 질의 건수 ==="
|
||||
kubectl -n "$NS" logs deploy/postgres --tail=-1 2>/dev/null \
|
||||
| tail -n +$((LOG_BEFORE + 1)) \
|
||||
| grep -F "$SID" \
|
||||
| grep -oE "^[0-9-]+ [0-9:.]+ [A-Z]+ \[[0-9]+\] [0-9.]+" \
|
||||
| awk '{print $NF}' | sort | uniq -c \
|
||||
| sed -e "s/$K0_IP/[keycloak-0]/" -e "s/$K1_IP/[keycloak-1]/" -e 's/^/ /'
|
||||
+207
@@ -0,0 +1,207 @@
|
||||
#!/usr/bin/env bash
|
||||
# Experiment 0 — is a session created on one Keycloak node usable on the other?
|
||||
#
|
||||
# Forming a cluster is not the same as sharing session state. The Infinispan log
|
||||
# says "cluster view (2)", but that only proves the members found each other.
|
||||
#
|
||||
# Design notes, learned the hard way:
|
||||
#
|
||||
# * Every probe has a CONTROL. A result from the far node means nothing unless
|
||||
# the same call against the issuing node is also measured. The first version
|
||||
# of this script reported "403 on keycloak-1" as if it were a replication
|
||||
# failure; the issuing node returned 403 too, and the cause was a missing
|
||||
# openid scope. Measure both, always.
|
||||
#
|
||||
# * Sessions are tracked by SID, not by count. Both the test login and the
|
||||
# admin API calls create sessions for the same user, so counts are noisy.
|
||||
# A specific session id either appears in a node's answer or it does not.
|
||||
#
|
||||
# * The probe is the REFRESH TOKEN grant, not userinfo. userinfo only validates
|
||||
# a signature and can succeed on a node that knows nothing about the session.
|
||||
# Refreshing requires the node to find the session, check it is alive, and
|
||||
# write back a new refresh time — it actually touches the session store.
|
||||
#
|
||||
# Talks to pod IPs directly: going through nginx/Traefik would hide which node
|
||||
# handled each request, which is the entire question.
|
||||
#
|
||||
# ./deploy/lab/scripts/experiment-session-replication.sh
|
||||
set -uo pipefail
|
||||
|
||||
NS="${NS:-keycloak-lab}"
|
||||
OUT="${OUT:-/tmp/session-replication}"
|
||||
mkdir -p "$OUT"
|
||||
|
||||
PSQL="kubectl -n $NS exec deploy/postgres -- psql -U keycloak -d keycloak -tAc"
|
||||
|
||||
echo "수집 시각: $(date '+%Y-%m-%d %H:%M:%S %Z')"
|
||||
echo
|
||||
|
||||
K0_IP=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.status.podIP}')
|
||||
K1_IP=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.status.podIP}')
|
||||
K0_NODE=$(kubectl -n "$NS" get pod keycloak-0 -o jsonpath='{.spec.nodeName}')
|
||||
K1_NODE=$(kubectl -n "$NS" get pod keycloak-1 -o jsonpath='{.spec.nodeName}')
|
||||
ADMIN_PW=$(kubectl -n "$NS" get secret keycloak-lab-secrets \
|
||||
-o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d)
|
||||
|
||||
echo "=== 대상 ==="
|
||||
printf ' keycloak-0 %-14s %s\n' "$K0_IP" "$K0_NODE"
|
||||
printf ' keycloak-1 %-14s %s\n' "$K1_IP" "$K1_NODE"
|
||||
echo
|
||||
|
||||
echo "=== [0] 실험 전 DB 세션 ==="
|
||||
$PSQL "select offline_flag, count(*) from offline_user_session group by offline_flag" 2>/dev/null \
|
||||
| sed 's/^/ offline_flag=/' || echo " (없음)"
|
||||
echo
|
||||
|
||||
# 파드 하나 안에서 전 단계를 실행한다. 단계마다 파드를 새로 띄우면 토큰을
|
||||
# 단계 사이로 넘길 수 없다.
|
||||
kubectl -n "$NS" run kc-probe --rm -i --restart=Never \
|
||||
--image=curlimages/curl:8.11.1 --quiet --command -- sh -c "
|
||||
set -u
|
||||
K0='http://$K0_IP:8080'; K1='http://$K1_IP:8080'
|
||||
TOKEN_EP='/realms/master/protocol/openid-connect/token'
|
||||
jget() { sed -n \"s/.*\\\"\$1\\\":\\\"\\([^\\\"]*\\)\\\".*/\\1/p\"; }
|
||||
|
||||
# ── [1] keycloak-0 에서 로그인. 이 노드가 세션의 출생지다 ──────────────────
|
||||
LOGIN=\$(curl -s -X POST \"\$K0\$TOKEN_EP\" \
|
||||
-d grant_type=password -d client_id=admin-cli \
|
||||
-d username=admin -d 'password=$ADMIN_PW')
|
||||
echo '###STEP1_LOGIN'; echo \"\$LOGIN\"
|
||||
|
||||
AT=\$(echo \"\$LOGIN\" | jget access_token)
|
||||
RT=\$(echo \"\$LOGIN\" | jget refresh_token)
|
||||
|
||||
# ── [2] 관리 API 조회용 토큰. 세션 오염을 피하려고 따로 하나만 더 만든다 ──
|
||||
ADMTOK=\$(curl -s -X POST \"\$K0\$TOKEN_EP\" \
|
||||
-d grant_type=password -d client_id=admin-cli \
|
||||
-d username=admin -d 'password=$ADMIN_PW' | jget access_token)
|
||||
CID=\$(curl -s -H \"Authorization: Bearer \$ADMTOK\" \
|
||||
\"\$K0/admin/realms/master/clients?clientId=admin-cli\" | jget id | head -1)
|
||||
|
||||
# ── [3] 두 노드에 같은 질문을 한다: admin-cli 의 세션 목록 ────────────────
|
||||
echo '###STEP3_SESSIONS_K0'
|
||||
curl -s -H \"Authorization: Bearer \$ADMTOK\" \
|
||||
\"\$K0/admin/realms/master/clients/\$CID/user-sessions?max=100\"
|
||||
echo
|
||||
echo '###STEP3_SESSIONS_K1'
|
||||
curl -s -H \"Authorization: Bearer \$ADMTOK\" \
|
||||
\"\$K1/admin/realms/master/clients/\$CID/user-sessions?max=100\"
|
||||
echo
|
||||
|
||||
# ── [4] 대조군: keycloak-0 이 발급한 refresh token 을 keycloak-0 에 쓴다 ──
|
||||
# 먼저 반대편에 써야 하므로 여기서는 쓰지 않고, 순서를 [5] 뒤로 미룬다.
|
||||
# refresh token 은 회전(rotation)되므로 한 번 쓰면 옛 것이 무효가 된다.
|
||||
# 따라서 '반대편 먼저'가 유일하게 의미 있는 순서다.
|
||||
|
||||
# ── [5] 시험군: keycloak-0 이 발급한 refresh token 을 keycloak-1 에 쓴다 ──
|
||||
echo '###STEP5_REFRESH_ON_K1'
|
||||
curl -s -w '\nhttp_code=%{http_code}\n' -X POST \"\$K1\$TOKEN_EP\" \
|
||||
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\"
|
||||
|
||||
RT2=\$(curl -s -X POST \"\$K1\$TOKEN_EP\" \
|
||||
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT\" \
|
||||
| jget refresh_token)
|
||||
|
||||
# ── [6] 무효화가 반대 방향으로도 전파되는가 ───────────────────────────────
|
||||
# keycloak-1 에서 로그아웃시키고, keycloak-0 에서 갱신을 시도한다.
|
||||
echo '###STEP6_LOGOUT_VIA_K1'
|
||||
curl -s -o /dev/null -w 'http_code=%{http_code}\n' -X POST \"\$K1/realms/master/protocol/openid-connect/logout\" \
|
||||
-d client_id=admin-cli -d \"refresh_token=\$RT2\"
|
||||
|
||||
echo '###STEP7_REFRESH_ON_K0_AFTER_LOGOUT'
|
||||
curl -s -w '\nhttp_code=%{http_code}\n' -X POST \"\$K0\$TOKEN_EP\" \
|
||||
-d grant_type=refresh_token -d client_id=admin-cli -d \"refresh_token=\$RT2\"
|
||||
echo '###END'
|
||||
" > "$OUT/raw.txt" 2>&1
|
||||
|
||||
sed -i '/^pod .* deleted$/d' "$OUT/raw.txt"
|
||||
|
||||
python3 - "$OUT/raw.txt" <<'PY' | tee "$OUT/report.txt"
|
||||
import base64, json, sys
|
||||
|
||||
raw = open(sys.argv[1]).read()
|
||||
blocks, cur = {}, None
|
||||
for line in raw.splitlines():
|
||||
if line.startswith('###'):
|
||||
cur = line[3:]; blocks[cur] = []
|
||||
elif cur is not None:
|
||||
blocks[cur].append(line)
|
||||
get = lambda k: '\n'.join(blocks.get(k, [])).strip()
|
||||
|
||||
def j(s):
|
||||
try: return json.JSONDecoder().raw_decode(s.strip())[0]
|
||||
except Exception: return None
|
||||
|
||||
def claims(tok):
|
||||
p = tok.split('.')[1]; p += '=' * (-len(p) % 4)
|
||||
return json.loads(base64.urlsafe_b64decode(p))
|
||||
|
||||
login = j(get('STEP1_LOGIN'))
|
||||
if not login or 'access_token' not in login:
|
||||
print('로그인 실패:', get('STEP1_LOGIN')[:300]); sys.exit(1)
|
||||
|
||||
ac = claims(login['access_token'])
|
||||
rc = claims(login['refresh_token'])
|
||||
SID = ac['sid']
|
||||
print('=== [1] keycloak-0 에서 로그인 ===')
|
||||
print(f" sid {SID}")
|
||||
print(f" sub {ac.get('sub')}")
|
||||
print(f" iss {ac.get('iss')}")
|
||||
print(f" access 수명 {ac['exp']-ac['iat']}초")
|
||||
print(f" refresh 수명 {rc['exp']-rc['iat']}초 typ={rc.get('typ')}")
|
||||
print(f" refresh jti {rc.get('jti')}")
|
||||
|
||||
print()
|
||||
print('=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===')
|
||||
for step, who in (('STEP3_SESSIONS_K0', 'keycloak-0 (발급 노드)'),
|
||||
('STEP3_SESSIONS_K1', 'keycloak-1 (반대편)')):
|
||||
d = j(get(step))
|
||||
if d is None:
|
||||
print(f' {who:24} 파싱 실패: {get(step)[:120]}'); continue
|
||||
ids = [s.get('id') for s in d]
|
||||
mark = '보임 ✔' if SID in ids else '없음 ✘'
|
||||
print(f' {who:24} 세션 {len(ids)}개 중 대상 sid → {mark}')
|
||||
for s in d:
|
||||
if s.get('id') == SID:
|
||||
print(f" ipAddress={s.get('ipAddress')} start={s.get('start')} lastAccess={s.get('lastAccess')}")
|
||||
|
||||
def show(step, title, expect):
|
||||
print(); print(f'=== {title} ===')
|
||||
body = get(step)
|
||||
code = [l for l in body.splitlines() if l.startswith('http_code=')]
|
||||
code = code[0].split('=')[1] if code else '?'
|
||||
d = j(body)
|
||||
ok = '기대대로' if code == expect else f'기대({expect})와 다름'
|
||||
print(f' HTTP {code} ← {ok}')
|
||||
if d and 'access_token' in d:
|
||||
c = claims(d['access_token'])
|
||||
same = '동일 ✔' if c.get('sid') == SID else f"다름 ✘ ({c.get('sid')})"
|
||||
print(f' 새 토큰의 sid → {same}')
|
||||
elif d:
|
||||
print(f" error {d.get('error')}")
|
||||
print(f" error_description {d.get('error_description')}")
|
||||
|
||||
show('STEP5_REFRESH_ON_K1',
|
||||
'[5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용', '200')
|
||||
|
||||
print(); print('=== [6] keycloak-1 을 통해 로그아웃 ===')
|
||||
print(' ' + get('STEP6_LOGOUT_VIA_K1').strip())
|
||||
|
||||
show('STEP7_REFRESH_ON_K0_AFTER_LOGOUT',
|
||||
'[7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파)', '400')
|
||||
|
||||
open('/tmp/session-replication/sid.txt','w').write(SID)
|
||||
PY
|
||||
|
||||
SID=$(cat /tmp/session-replication/sid.txt 2>/dev/null)
|
||||
echo
|
||||
echo "=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ==="
|
||||
echo " 대상 sid: $SID"
|
||||
$PSQL "select user_session_id, offline_flag, created_on, last_session_refresh
|
||||
from offline_user_session where user_session_id='$SID'" 2>/dev/null \
|
||||
| sed 's/^/ /' | grep -q . \
|
||||
&& $PSQL "select user_session_id||' | flag='||offline_flag||' | created='||created_on||' | refresh='||last_session_refresh
|
||||
from offline_user_session where user_session_id='$SID'" 2>/dev/null | sed 's/^/ /' \
|
||||
|| echo " 행 없음 — 로그아웃으로 삭제되었다"
|
||||
echo
|
||||
echo " 전체 세션 수: $($PSQL 'select count(*) from offline_user_session' 2>/dev/null)"
|
||||
@@ -0,0 +1,13 @@
|
||||
=== [기준선 1] 클러스터 뷰 — 양쪽 파드의 마지막 ISPN000094 ===
|
||||
keycloak-0: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
|
||||
keycloak-1: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
|
||||
|
||||
=== [기준선 2] JGROUPS_PING — 디스커버리 등록 ===
|
||||
name | ip | coord | coordinated_by
|
||||
------------------+-----------------+-------+---------------------------------------------
|
||||
keycloak-0-30843 | 10.42.1.43:7800 | f | uuid://00000000-0000-0000-0000-000000000007
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t | uuid://00000000-0000-0000-0000-000000000007
|
||||
(2 rows)
|
||||
|
||||
=== [기준선 3] 기존 NetworkPolicy ===
|
||||
No resources found in keycloak-lab namespace.
|
||||
@@ -0,0 +1,17 @@
|
||||
=== [대조군] 차단 전 — keycloak-0 로그인 → keycloak-1 에서 refresh ===
|
||||
sid tAWs2gCPr6SOcD4jDR9-_CzB
|
||||
keycloak-1 에서 refresh: 200
|
||||
|
||||
=== [기준선 4] JGroups 지표 — 양쪽 노드 ===
|
||||
--- K0 (10.42.1.43) ---
|
||||
vendor_jgroups_stats_bytes_sent_total 31476.0
|
||||
vendor_jgroups_merge3_get_num_merge_events 0.0
|
||||
vendor_jgroups_merge3_get_views 0.0
|
||||
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
|
||||
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
|
||||
--- K1 (10.42.0.35) ---
|
||||
vendor_jgroups_merge3_get_views 0.0
|
||||
vendor_jgroups_stats_bytes_sent_total 126765.0
|
||||
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
|
||||
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
|
||||
vendor_jgroups_merge3_get_num_merge_events 0.0
|
||||
@@ -0,0 +1,8 @@
|
||||
=== 차단 적용 ===
|
||||
networkpolicy.networking.k8s.io/a1-block-jgroups-transport created
|
||||
a1-block-jgroups-transport map[app:keycloak]
|
||||
|
||||
적용 시각: 11:38:08
|
||||
=== FD_SOCK2 가 상대를 의심하기까지 기다린다 (15초 간격, 최대 3분) ===
|
||||
+15초 suspected(k0 k1) =
|
||||
→ 변화 감지
|
||||
@@ -0,0 +1,26 @@
|
||||
차단 경과: 11:38:51 (적용 11:38:08)
|
||||
|
||||
=== [차단 후 1] JGroups 지표 ===
|
||||
--- keycloak-0 ---
|
||||
vendor_jgroups_stats_bytes_sent_total 32857.0
|
||||
vendor_jgroups_merge3_get_num_merge_events 0.0
|
||||
vendor_jgroups_merge3_get_views 0.0
|
||||
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
|
||||
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
|
||||
--- keycloak-1 ---
|
||||
vendor_jgroups_merge3_get_views 0.0
|
||||
vendor_jgroups_stats_bytes_sent_total 129103.0
|
||||
vendor_jgroups_nakack2_get_xmit_table_missing_messages 0.0
|
||||
vendor_jgroups_fd_sock2_get_num_suspected_members 0.0
|
||||
vendor_jgroups_merge3_get_num_merge_events 0.0
|
||||
|
||||
=== [차단 후 2] 클러스터 뷰 — 갈라졌는가 ===
|
||||
keycloak-0:
|
||||
keycloak-1:
|
||||
=== [차단 후 3] JGROUPS_PING — 디스커버리는 살아 있는가 ===
|
||||
name | ip | coord
|
||||
------------------+-----------------+-------
|
||||
keycloak-0-30843 | 10.42.1.43:7800 | f
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t
|
||||
(2 rows)
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
=== [문제 확정] NetworkPolicy 적용 후에도 기존 연결이 conntrack 에 살아 있다 ===
|
||||
--- kc-lab-1 ---
|
||||
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1
|
||||
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 src=10.42.1.43 dst=10.42.0.35 sport=57800 dport=50477 [ASSURED] mark=0 use=1
|
||||
--- kc-lab-2 ---
|
||||
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1
|
||||
tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED] src=10.42.0.35 dst=10.42.1.58 sport=7800 dport=34824 mark=0 use=1
|
||||
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 src=10.42.1.43 dst=10.42.0.35 sport=57800 dport=50477 [ASSURED] mark=0 use=1
|
||||
|
||||
=== [조치] 7800 흐름의 conntrack 항목을 지운다 → 다음 패킷이 정책을 다시 탄다 ===
|
||||
kc-lab-1: tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 src=10.42.1.43 dst=10.42.0.35 sport=7800 dport=40023 [ASSURED] mark=0 use=1 conntrack v1.4.7 (conntrack-tools): 0 flow entries have been deleted.
|
||||
kc-lab-2: tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED] src=10.42.0.35 dst=10.42.1.58 sport=7800 dport=34824 mark=0 use=1 conntrack v1.4.7 (conntrack-tools): 0 flow entries have been deleted.
|
||||
|
||||
=== 삭제 후 7800 conntrack ===
|
||||
kc-lab-1: 2 건
|
||||
kc-lab-2: 2 건
|
||||
@@ -0,0 +1,13 @@
|
||||
관찰 시작: 11:42:03
|
||||
+20초 suspected(k0 k1) = []
|
||||
+40초 suspected(k0 k1) = []
|
||||
+60초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
|
||||
+80초 suspected(k0 k1) = []
|
||||
+100초 suspected(k0 k1) = []
|
||||
+120초 suspected(k0 k1) = []
|
||||
+140초 suspected(k0 k1) = [0.0 ]
|
||||
+160초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
|
||||
|
||||
=== 클러스터 뷰 변화 (최근 8분) ===
|
||||
--- keycloak-0 ---
|
||||
--- keycloak-1 ---
|
||||
@@ -0,0 +1,16 @@
|
||||
=== vendor_cluster_size — 지난 25분 (차단 11:38:08, conntrack 삭제 11:41) ===
|
||||
keycloak-0:
|
||||
11:20=2 11:21=2 11:22=2 11:23=2 11:24=2 11:25=2 11:26=2 11:27=2 11:28=2 11:29=2 11:30=2 11:31=2 11:32=2 11:33=2 11:34=2 11:35=2 11:36=2 11:37=2 11:38=2 11:39=2 11:40=2 11:41=2 11:42=2 11:43=2 11:44=2 11:45=2
|
||||
keycloak-1:
|
||||
11:20=2 11:21=2 11:22=2 11:23=2 11:24=2 11:25=2 11:26=2 11:27=2 11:28=2 11:29=2 11:30=2 11:31=2 11:32=2 11:33=2 11:34=2 11:35=2 11:36=2 11:37=2 11:38=2 11:39=2 11:40=2 11:41=2 11:42=2 11:43=2 11:44=2 11:45=2
|
||||
|
||||
=== 현재 값 ===
|
||||
keycloak-1 = 2 멤버
|
||||
keycloak-0 = 2 멤버
|
||||
|
||||
=== 7800 소켓 상태 (파드 내부) ===
|
||||
keycloak-0 2
|
||||
keycloak-1 2
|
||||
=== conntrack ===
|
||||
kc-lab-1 1 건
|
||||
kc-lab-2 1 건
|
||||
@@ -0,0 +1,9 @@
|
||||
=== 정책이 걸린 상태에서 keycloak-0 을 재시작한다 → 재연결이 막힌다 ===
|
||||
재시작 시각: 11:46:07
|
||||
pod "keycloak-0" deleted from keycloak-lab namespace
|
||||
keycloak-0 false 10.42.1.67 2026-09-04T02:44:23Z
|
||||
|
||||
=== cluster_size 추이 ===
|
||||
keycloak-0: 11:45:27=1 11:45:57=1 11:46:27=1 11:46:57=1 11:47:27=1
|
||||
keycloak-0: 11:40:57=2 11:41:27=2 11:41:57=2 11:42:27=2 11:42:57=2 11:43:27=2 11:43:57=2
|
||||
keycloak-1: 11:40:57=2 11:41:27=2 11:41:57=2 11:42:27=2 11:42:57=2 11:43:27=2 11:43:57=2 11:44:27=1 11:44:57=1 11:45:27=1 11:45:57=1 11:46:27=1 11:46:57=1 11:47:27=1
|
||||
@@ -0,0 +1,16 @@
|
||||
=== keycloak-0 헬스 상태 ===
|
||||
keycloak-0 = 10.42.1.67 keycloak-1 = 10.42.0.35
|
||||
PodReadyToStartContainers=True
|
||||
Initialized=True
|
||||
Ready=False ContainersNotReady
|
||||
ContainersReady=False ContainersNotReady
|
||||
PodScheduled=True
|
||||
|
||||
=== ★ 본 시험 — 분단 상태에서 교차 노드 세션이 되는가 ===
|
||||
[1] keycloak-0 로그인 sid=nShl5TaBrZnKStDqaspjgmJB
|
||||
[2] keycloak-1 에서 refresh HTTP 200
|
||||
[3] keycloak-1 에서 로그아웃 HTTP 204
|
||||
[4] keycloak-0 에서 재갱신 시도 HTTP 200
|
||||
(400 이면 무효화가 전파된 것)
|
||||
|
||||
=== DB 세션 수 ===
|
||||
@@ -0,0 +1,33 @@
|
||||
=== 그 sid 가 DB 에 남아 있는가 ===
|
||||
user_session_id | offline_flag | last_session_refresh
|
||||
-----------------+--------------+----------------------
|
||||
(0 rows)
|
||||
|
||||
=== 전체 온라인 세션 수 ===
|
||||
1
|
||||
|
||||
=== 노드별 세션 캐시 엔트리 (Prometheus) ===
|
||||
keycloak-1 kc-lab-1 = 0
|
||||
keycloak-0 kc-lab-2 = 1
|
||||
|
||||
=== keycloak-0 이 Ready 가 아닌 이유 — 헬스 응답 ===
|
||||
{
|
||||
"status": "DOWN",
|
||||
"checks": [
|
||||
{
|
||||
"name": "Graceful Shutdown",
|
||||
"status": "UP"
|
||||
},
|
||||
{
|
||||
"name": "Keycloak cluster health check",
|
||||
"status": "DOWN",
|
||||
"data": {
|
||||
"Failing since": "2026-09-04 02:45:14,251"
|
||||
}
|
||||
},
|
||||
{
|
||||
"name": "Keycloak database connections async health check",
|
||||
"status": "UP"
|
||||
},
|
||||
{
|
||||
"name": "Keycloak Initialized",
|
||||
@@ -0,0 +1,25 @@
|
||||
=== 양쪽 노드의 readiness — 둘 다 DOWN 이면 전면 장애다 ===
|
||||
Traceback (most recent call last):
|
||||
File "<string>", line 3, in <module>
|
||||
d=json.load(sys.stdin)
|
||||
File "/usr/lib/python3.14/json/__init__.py", line 298, in load
|
||||
return loads(fp.read(),
|
||||
cls=cls, object_hook=object_hook,
|
||||
parse_float=parse_float, parse_int=parse_int,
|
||||
parse_constant=parse_constant, object_pairs_hook=object_pairs_hook, **kw)
|
||||
File "/usr/lib/python3.14/json/__init__.py", line 352, in loads
|
||||
return _default_decoder.decode(s)
|
||||
~~~~~~~~~~~~~~~~~~~~~~~^^^
|
||||
File "/usr/lib/python3.14/json/decoder.py", line 348, in decode
|
||||
raise JSONDecodeError("Extra data", s, end)
|
||||
json.decoder.JSONDecodeError: Extra data: line 21 column 2 (char 446)
|
||||
|
||||
=== 파드 Ready 상태 ===
|
||||
keycloak-0 false 0
|
||||
keycloak-1 true 0
|
||||
|
||||
=== ★ Service 엔드포인트 — 트래픽을 받는 파드가 남아 있는가 ===
|
||||
ready 주소: [10.42.0.35] notReady : [10.42.1.67]
|
||||
=== ★ 외부 진입점으로 실제 로그인이 되는가 (nginx→Traefik→Service) ===
|
||||
https://auth.hyeonworks.com/realms/master HTTP 200
|
||||
토큰 발급 HTTP 200
|
||||
@@ -0,0 +1,24 @@
|
||||
=== 차단 해제 ===
|
||||
해제 시각: 11:49:58
|
||||
networkpolicy.networking.k8s.io "a1-block-jgroups-transport" deleted from keycloak-lab namespace
|
||||
|
||||
=== 자동으로 다시 붙는가 (30초 간격, 최대 4분) ===
|
||||
+30초 keycloak-0=1 keycloak-1=1 | Ready 파드 2 개
|
||||
+60초 keycloak-0=1 keycloak-1=1 | Ready 파드 2 개
|
||||
+90초 keycloak-0=2 keycloak-1=2 | Ready 파드 3 개
|
||||
→ 클러스터 재형성
|
||||
|
||||
=== 복구 로그 ===
|
||||
keycloak-0: [keycloak-0-26403(v=16.0.12)|0] (1) [keycloak-0-26403(v=16.0.12)]
|
||||
keycloak-1: [keycloak-1-48749(v=16.0.12)|6] (1) [keycloak-1-48749(v=16.0.12)]
|
||||
|
||||
=== MERGE3 가 합쳤는가 ===
|
||||
merge_events keycloak-1 = 1
|
||||
merge_events keycloak-0 = 1
|
||||
=== JGROUPS_PING — 코디네이터가 하나로 돌아왔는가 ===
|
||||
name | ip | coord
|
||||
------------------+-----------------+-------
|
||||
keycloak-0-26403 | 10.42.1.67:7800 | t
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | f
|
||||
(2 rows)
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
# A-1 — JGroups 트랜스포트(7800) 차단 증거
|
||||
|
||||
2026-09-04 11:38–11:52 KST · Keycloak 26.7.0 / Infinispan 16.0.12
|
||||
해설: [`docs/experiment-a1-jgroups-transport-block.md`](../../experiment-a1-jgroups-transport-block.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-baseline-cluster.txt` | 차단 전 — 양쪽이 뷰 ID 5·멤버 2로 일치, `JGROUPS_PING` 코디네이터 1명 |
|
||||
| `02-control-before-block.txt` | **대조군** — 차단 전 교차 노드 refresh `200`, JGroups 지표 전부 0 |
|
||||
| `03-block-applied.txt` | NetworkPolicy 적용. **빈 측정값을 "변화 감지"로 오판한 기록** |
|
||||
| `04-after-block-state.txt` | 차단 43초 후 — 지표 무변화, `JGROUPS_PING` 그대로 |
|
||||
| `05-conntrack-problem.txt` | **핵심 문제** — `ESTABLISHED [ASSURED]` 로 기존 연결이 살아 있음. FD_SOCK2 의 **57800** 포트도 함께 드러남 |
|
||||
| `06-partition-observed.txt` | 임시 curl 파드 폴링의 실패 — 빈 값·개수 불일치 |
|
||||
| `07-cluster-size.txt` | **`vendor_cluster_size` 가 25분 내내 2** — 분단이 일어나지 않았다는 결정적 증거 |
|
||||
| `08-restart-forced-partition.txt` | 재연결 강제 후 `2 → 1` |
|
||||
| `09-cross-node-under-partition.txt` | **본 시험** — 교차 refresh `200`(예측 적중), **로그아웃 후 재갱신 `200`(예측 빗나감)** |
|
||||
| `10-logout-not-propagated.txt` | 기제 확정 — **DB 행 0건인데 keycloak-0 캐시에 1건**, 헬스체크 `cluster health: DOWN` |
|
||||
| `11-service-impact.txt` | **분단 노드가 Service 에서 빠짐.** `ready=[10.42.0.35] notReady=[10.42.1.67]`, 외부 로그인 `200` |
|
||||
| `12-recovery.txt` | 90초 만에 자동 재형성, `merge3_get_num_merge_events = 1`, 코디네이터 재선출 |
|
||||
| `a1-cluster-size-partition-recovery.png` | Grafana — `vendor_cluster_size` 가 `2 → 1 → 2` 로 움직이는 전 구간 |
|
||||
|
||||
## 핵심 세 줄
|
||||
|
||||
1. **NetworkPolicy 만으로는 이미 붙어 있는 클러스터를 못 끊는다.** conntrack 의 ESTABLISHED 가 먼저 통과시킨다.
|
||||
2. **세션 공유는 분단을 견딘다(200).** 통념이 틀렸고 A-0 모델이 맞다.
|
||||
3. **로그아웃 무효화는 7800 을 탄다.** DB 행이 지워져도 반대편은 낡은 캐시로 200 을 준다 — A-0 의 인과 해석을 정정한다.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 66 KiB |
@@ -0,0 +1,14 @@
|
||||
=== A-2 기준선 — 클러스터가 정상으로 돌아왔는가 ===
|
||||
keycloak-0 true 10.42.1.67 kc-lab-2
|
||||
keycloak-1 true 10.42.0.35 kc-lab-1
|
||||
postgres-7b474b88c8-sn9ff true 10.42.1.24 kc-lab-2
|
||||
|
||||
cluster_size keycloak-1 = 2
|
||||
cluster_size keycloak-0 = 2
|
||||
|
||||
=== 노드별 세션 캐시 (실험 설계에 필요) ===
|
||||
keycloak-1 kc-lab-1 = 0 건
|
||||
keycloak-0 kc-lab-2 = 0 건
|
||||
|
||||
=== DB 온라인 세션 ===
|
||||
2
|
||||
@@ -0,0 +1,10 @@
|
||||
pod/a2-probe condition met
|
||||
keycloak-0=10.42.1.67 keycloak-1=10.42.0.35
|
||||
|
||||
=== [준비] 양쪽 노드에 세션을 하나씩 만든다 ===
|
||||
keycloak-0 에서 로그인 sid=EAXV5HcG2J1BZ3vnwONf64AQ 토큰길이=613
|
||||
keycloak-1 에서 로그인 sid=McyTj5lj3n_JqApCXeuAHExc 토큰길이=613
|
||||
|
||||
=== [확인] 세션이 각자 노드에만 캐시되었는가 ===
|
||||
keycloak-1 = 0 건
|
||||
keycloak-0 = 1 건
|
||||
@@ -0,0 +1,16 @@
|
||||
=== [1] 토큰을 새로 발급 (access 수명 60초) ===
|
||||
발급 완료 sid=RKXQGAgkuLtouFMVPTFmp_0_
|
||||
|
||||
=== [2] PostgreSQL 정지 ===
|
||||
정지 시각: 11:56:04
|
||||
deployment.apps/postgres scaled
|
||||
pod/postgres-7b474b88c8-sn9ff condition met
|
||||
삭제 완료: 11:56:04
|
||||
|
||||
=== [3] 네 경로를 즉시 시험 ===
|
||||
④ 이미 발급된 access token 으로 관리 API HTTP 000000{"error":"HTTP 401 Unauthorized"}401
|
||||
① 캐시를 가진 노드(keycloak-0)에서 refresh HTTP 500
|
||||
② 캐시가 없는 노드(keycloak-1)에서 refresh HTTP 500
|
||||
③ 새 로그인 HTTP 500
|
||||
--- 오류 본문 (새 로그인) ---
|
||||
{"error":"unknown_error","error_description":"For more on this error consult the server log."}
|
||||
@@ -0,0 +1,29 @@
|
||||
=== 파드 Ready 상태 — DB 가 없으면 어떻게 되는가 ===
|
||||
keycloak-0 false 0
|
||||
keycloak-1 false 0
|
||||
|
||||
=== Service 엔드포인트 ===
|
||||
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
|
||||
Warning: v1 Endpoints is deprecated in v1.33+; use discovery.k8s.io/v1 EndpointSlice
|
||||
notReady: [10.42.0.35 10.42.1.67]
|
||||
=== health/ready 상세 ===
|
||||
전체: DOWN
|
||||
Graceful Shutdown UP
|
||||
Keycloak cluster health check UP
|
||||
Keycloak database connections async health check DOWN
|
||||
Keycloak Initialized UP
|
||||
|
||||
=== ④ 다시 — 서명 검증만 필요한 경로는 살아 있는가 ===
|
||||
JWKS 엔드포인트(realm 공개키) HTTP 200
|
||||
realm 메타데이터(.well-known) HTTP 200
|
||||
관리 API(세션 조회 필요) HTTP 500
|
||||
|
||||
=== 외부 진입점 ===
|
||||
https://auth.hyeonworks.com/realms/master HTTP 503
|
||||
|
||||
=== Keycloak 로그 — 실제 오류 ===
|
||||
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call(ConnectionPool.java:664)
|
||||
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call(ConnectionPool.java:645)
|
||||
Caused by: java.net.ConnectException: Connection refused
|
||||
at org.postgresql.core.v3.ConnectionFactoryImpl.tryConnect(ConnectionFactoryImpl.java:219)
|
||||
at org.postgresql.core.v3.ConnectionFactoryImpl.openConnectionImpl(ConnectionFactoryImpl.java:365)
|
||||
@@ -0,0 +1,21 @@
|
||||
=== ★ up 지표는 무엇을 말하는가 (프로세스는 살아 있다) ===
|
||||
up{pod=keycloak-1} = 1 ← 1 인데 서비스는 503 이다
|
||||
up{pod=keycloak-0} = 1 ← 1 인데 서비스는 503 이다
|
||||
|
||||
=== 복구 — PostgreSQL 재기동 ===
|
||||
재기동 시각: 11:57:09
|
||||
deployment.apps/postgres scaled
|
||||
Waiting for deployment "postgres" rollout to finish: 0 out of 1 new replicas have been updated...
|
||||
Waiting for deployment "postgres" rollout to finish: 0 of 1 updated replicas are available...
|
||||
deployment "postgres" successfully rolled out
|
||||
|
||||
=== Keycloak 이 스스로 회복하는가 (재시작 없이) ===
|
||||
+15초 keycloak-0 true keycloak-1 true | 외부 HTTP 200
|
||||
→ 서비스 복귀
|
||||
|
||||
=== 재시작 횟수 — 파드가 죽었다 살아난 것인가, 그대로 회복한 것인가 ===
|
||||
keycloak-0 0
|
||||
keycloak-1 0
|
||||
|
||||
=== 정지 전 세션이 살아남았는가 ===
|
||||
online 세션 5
|
||||
@@ -0,0 +1,19 @@
|
||||
# A-2 — PostgreSQL 정지 증거
|
||||
|
||||
2026-09-04 11:56–11:58 KST · Keycloak 26.7.0
|
||||
해설: [`docs/experiment-a2-database-loss.md`](../../experiment-a2-database-loss.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-baseline.txt` | 정지 전 — 양쪽 Ready, `cluster_size=2` |
|
||||
| `02-setup-sessions.txt` | 양쪽 노드에 세션 하나씩. 캐시는 각자 노드에만 |
|
||||
| `03-four-paths.txt` | **네 경로 전부 `500`.** 캐시를 가진 노드도 실패 — refresh 는 쓰기다 |
|
||||
| `04-health-and-service.txt` | **전면 장애 증거** — Ready 파드 0개, `ready 주소=[]`, 외부 **503**, `database connections: DOWN`. JWKS·.well-known 은 `200` |
|
||||
| `05-recovery.txt` | **`up=1` 인 채로 503.** DB 복귀 15초 후 재시작 0회로 자동 회복, 세션 5건 생존 |
|
||||
| `a2-up-stayed-1-during-outage.png` | Grafana — `up{job="keycloak"}` 이 전면 장애 내내 **1에 평평** |
|
||||
|
||||
## 핵심 세 줄
|
||||
|
||||
1. **DB 는 단일 장애점이다.** Keycloak 을 몇 대로 늘려도 같이 죽는다 — Ready 파드 0개, 외부 503.
|
||||
2. **캐시는 읽기를 대신할 뿐 쓰기를 못 한다.** refresh 는 `UPDATE LAST_SESSION_REFRESH` 를 하므로 캐시가 있어도 실패한다.
|
||||
3. **`up` 은 이 장애를 못 잡는다.** 알림은 readiness 와 외부 응답 코드에 걸어야 한다.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 60 KiB |
@@ -0,0 +1,19 @@
|
||||
=== [준비] 손실 측정 설계 확인 ===
|
||||
LAST_SESSION_REFRESH 는 integer(초) — 200ms 손실은 보이지 않는다
|
||||
created_on | integer | | not null |
|
||||
last_session_refresh | integer | | not null | 0
|
||||
"idx_user_session_expiration_created" btree (realm_id, offline_flag, remember_me, created_on, user_session_id, user_id)
|
||||
"idx_user_session_expiration_last_refresh" btree (realm_id, offline_flag, remember_me, last_session_refresh, user_session_id, user_id)
|
||||
→ 대신 행 존재 여부로 잰다. 로그인 하나 = 행 하나 = 이진 판정
|
||||
|
||||
전역 synchronous_commit: on
|
||||
|
||||
=== [1] 빠른 연속 로그인을 백그라운드로 시작 ===
|
||||
루프 시작
|
||||
6초 경과 — 지금까지 성공한 로그인: 0
|
||||
|
||||
=== [2] PostgreSQL 강제 종료 (SIGKILL) ===
|
||||
종료 시각: 12:00:26.511
|
||||
pod "postgres-7b474b88c8-xc2vt" force deleted from keycloak-lab namespace
|
||||
삭제 반환: 12:00:26.586
|
||||
클라이언트가 200 을 받은 로그인 수: 0
|
||||
@@ -0,0 +1,14 @@
|
||||
deployment "postgres" successfully rolled out
|
||||
|
||||
=== crash recovery 가 실행되었는가 (강제 종료의 흔적) ===
|
||||
2026-09-04 02:58:41.036 UTC [1] LOG: database system is ready to accept connections
|
||||
|
||||
=== [설계 확인] 로그인 트랜잭션도 synchronous_commit 을 끄는가 ===
|
||||
--- 로그인 트랜잭션 (INSERT 가 있는 것) ---
|
||||
2:BEGIN
|
||||
5:COMMIT
|
||||
6:BEGIN
|
||||
9:insert into OFFLINE_USER_SESSION (BROKER_SESSION_ID,CREATED_ON,DATA,LAST_SESSION_REFRESH,REALM_ID,REMEMBER_ME,USER_ID,VERSION,OFFLINE_FLAG,USER_SESSION_ID) values ($1,$2,$3,$4,$5,$6,$7,$8,$9,$10)
|
||||
10:insert into OFFLINE_CLIENT_SESSION (DATA,REALM_ID,TIMESTAMP,VERSION,CLIENT_ID,CLIENT_STORAGE_PROVIDER,EXTERNAL_CLIENT_ID,OFFLINE_FLAG,USER_SESSION_ID) values ($1,$2,$3,$4,$5,$6,$7,$8,$9)
|
||||
11:SET LOCAL synchronous_commit TO OFF
|
||||
12:COMMIT
|
||||
@@ -0,0 +1,17 @@
|
||||
=== [1] 로그인 루프 시작 (호스트에서 백그라운드로 exec — 세션이 살아 있어야 한다) ===
|
||||
8초 동안 클라이언트가 200 을 받은 로그인: 106 건
|
||||
|
||||
=== [2] SIGKILL ===
|
||||
종료: 12:01:32.981
|
||||
반환: 12:01:33.236
|
||||
최종 성공 로그인 수: 110 건
|
||||
마지막 sid: FimM-krSybBACP2qIvshLWwU
|
||||
마지막 sid: EwFfFwOIfqiv8N5GQ5OjtsVq
|
||||
마지막 sid: CJX-PxFkS7rUc_9FQgB7iw1f
|
||||
마지막 sid: 1EFK7SgUA4M7tq_SkC_BD2er
|
||||
마지막 sid: _LiqTczuyxlpOs3T3xs25SLv
|
||||
|
||||
=== [3] PostgreSQL 재기동 후 crash recovery 확인 ===
|
||||
Waiting for deployment "postgres" rollout to finish: 0 of 1 updated replicas are available...
|
||||
deployment "postgres" successfully rolled out
|
||||
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
|
||||
@@ -0,0 +1,27 @@
|
||||
=== [4] 클라이언트가 받은 sid 가 DB 에 있는가 ===
|
||||
클라이언트가 200 을 받은 sid: 291 건
|
||||
DB 온라인 세션 총계: 375
|
||||
|
||||
--- 마지막 15건을 하나씩 조회 ---
|
||||
TCCOYnVlN30Y2fGyJEsVJ_Lq 있음
|
||||
vBcllIkKWovh-FXN9tSzmxAe 있음
|
||||
eMpN_ywUdRTks9uBE9aTumPK 있음
|
||||
aPC_T0yrlMjrlskcLAp0AvY6 있음
|
||||
UBPxmduB-ahGHBg636sY3AEz 있음
|
||||
HLnBloNX9R3qQJSkpOnkNqL4 있음
|
||||
_KPJS30IAqVhkTJGHcCxKvrM 있음
|
||||
rq3caZ9MkyMYlFSSzRjQLygD 있음
|
||||
ivvQm70hjl55DPpF7vpYmz_E 있음
|
||||
81mx-rmi-tAeogHx3-su3z2q 있음
|
||||
WnvNDH93uzcz1XNFbaMSXk1A 있음
|
||||
DXPAIhjO5sGpCUlcD8IEoS8L 있음
|
||||
aZMvl4IwPdK-rC_7bG005Z5C 있음
|
||||
eIuBCprfWA5x0glcgSKYrrX0 있음
|
||||
ozES5kEeu2IFf_cfcC_jFlbF 있음
|
||||
|
||||
마지막 15건 중 유실: 0 건
|
||||
|
||||
=== [5] 전체 대조 — 몇 건이나 사라졌는가 ===
|
||||
클라이언트 성공: 291 건
|
||||
DB 에 존재: 291 건
|
||||
★ 유실: 0 건
|
||||
@@ -0,0 +1,12 @@
|
||||
=== [정리] 세션 테이블 비우고 루프 잔여 확인 ===
|
||||
DELETE 375
|
||||
남은 세션: 0
|
||||
|
||||
=== [재주입] postmaster(PID 1)에 SIGKILL — 진짜 크래시 ===
|
||||
8초 후 성공 로그인: 110 건
|
||||
SIGKILL: 12:03:21.441
|
||||
최종 성공 로그인: 139 건
|
||||
|
||||
=== [검증] 이번엔 crash recovery 가 돌았는가 ===
|
||||
deployment "postgres" successfully rolled out
|
||||
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
|
||||
@@ -0,0 +1,17 @@
|
||||
=== 로그인 루프 시작 ===
|
||||
8초 후: 112 건
|
||||
|
||||
=== 백엔드 프로세스에 SIGKILL → postmaster 가 재초기화한다 ===
|
||||
시각: 12:04:22.063
|
||||
최종 성공 로그인: 153 건
|
||||
|
||||
=== [검증] crash recovery 가 돌았는가 ===
|
||||
2026-09-04 02:59:48.427 UTC [1] LOG: database system is ready to accept connections
|
||||
2026-09-04 03:02:35.807 UTC [1] LOG: server process (PID 40) was terminated by signal 9: Killed
|
||||
2026-09-04 03:02:35.807 UTC [1] LOG: terminating any other active server processes
|
||||
2026-09-04 03:02:35.814 UTC [1] LOG: all server processes terminated; reinitializing
|
||||
2026-09-04 03:02:35.896 UTC [2585] LOG: database system was not properly shut down; automatic recovery in progress
|
||||
2026-09-04 03:02:35.899 UTC [2585] LOG: redo starts at 0/23CAB68
|
||||
2026-09-04 03:02:35.904 UTC [2585] LOG: redo done at 0/2529E40 system usage: CPU: user: 0.00 s, system: 0.00 s, elapsed: 0.00 s
|
||||
2026-09-04 03:02:35.923 UTC [2586] LOG: checkpoint complete: wrote 113 buffers (0.7%); 0 WAL file(s) added, 0 removed, 0 recycled; write=0.004 s, sync=0.004 s, total=0.015 s; sync files=27, longest=0.003 s, average=0.001 s; distance=1405 kB, estimate=1405 kB; lsn=0/252A048, redo lsn=0/252A048
|
||||
2026-09-04 03:02:35.926 UTC [1] LOG: database system is ready to accept connections
|
||||
@@ -0,0 +1,19 @@
|
||||
=== 크래시 전후 대조 ===
|
||||
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
|
||||
그중 DB 에 실제로 존재 : 149 건
|
||||
★ 유실 : 4 건
|
||||
DB 전체 온라인 세션 : 150 건
|
||||
|
||||
=== 유실된 sid 목록 ===
|
||||
★ CQUfg9HLH29xvhiu6pVlfWOo ← 토큰은 발급됐는데 세션이 없다
|
||||
★ 5gLP4fqmpZBbjhH_d-0TPMMr ← 토큰은 발급됐는데 세션이 없다
|
||||
★ hkcOv1QskUFmYveMLB6Hljra ← 토큰은 발급됐는데 세션이 없다
|
||||
★ p5XybeQIYmAs818gO4Vl_5ea ← 토큰은 발급됐는데 세션이 없다
|
||||
|
||||
=== 그 토큰이 지금 실제로 쓰이는가 (마지막 sid 로 확인) ===
|
||||
마지막 sid: 8do0Bw6tkVLDVxgxotE7GosH
|
||||
user_session_id | created_on | last_session_refresh
|
||||
--------------------------+------------+----------------------
|
||||
8do0Bw6tkVLDVxgxotE7GosH | 1788490958 | 1788490958
|
||||
(1 row)
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
# A-3 — DB 강제 종료와 데이터 손실 증거
|
||||
|
||||
2026-09-04 12:00–12:05 KST · Keycloak 26.7.0 / PostgreSQL 16
|
||||
해설: [`docs/experiment-a3-database-crash.md`](../../experiment-a3-database-crash.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-crash-injection.txt` | 첫 시도 실패 — 파드 안 백그라운드 루프가 `exec` 종료와 함께 죽어 0건 수집 |
|
||||
| `02-design-check.txt` | **핵심 설계 확인** — 로그인 트랜잭션도 `SET LOCAL synchronous_commit TO OFF` 로 커밋한다 |
|
||||
| `03-loss-measurement.txt` | `--grace-period=0 --force` 주입 |
|
||||
| `04-comparison.txt` | **유실 0건** — 그러나 crash recovery 가 안 돌았다. 죽인 적이 없는 것 |
|
||||
| `05-true-crash.txt` | `kill -9 1` 시도 — **컨테이너 안에서 PID 1 은 SIGKILL 을 무시한다** |
|
||||
| `06-backend-kill-crash.txt` | **성공한 주입** — 백엔드에 SIGKILL → `not properly shut down` / `redo starts` / `redo done` |
|
||||
| `07-loss-result.txt` | **결과: 153건 중 4건 유실.** 토큰은 발급됐는데 세션 행이 없는 sid 목록 |
|
||||
|
||||
## 핵심 세 줄
|
||||
|
||||
1. **로그인도 비동기 커밋이다.** refresh 시각뿐 아니라 **로그인 자체**가 사라질 수 있다.
|
||||
2. **153건 중 4건(약 2.6%) 유실** — 초당 19건 기준 마지막 0.2초 분량, `wal_writer_delay` 기본값과 일치.
|
||||
3. **주입을 세 번 시도해 세 번째에 성공했다.** 앞의 둘은 "손실 0"으로 보였지만 실제로는 크래시가 아니었다.
|
||||
@@ -0,0 +1,20 @@
|
||||
=== A-4 기준선 ===
|
||||
kc-lab-1 Ready true
|
||||
kc-lab-2 Ready <none>
|
||||
|
||||
a2-probe true kc-lab-2
|
||||
keycloak-0 true kc-lab-2
|
||||
keycloak-1 true kc-lab-1
|
||||
postgres-7b474b88c8-2gf27 true kc-lab-2
|
||||
|
||||
=== PVC 가 어느 노드에 묶여 있는가 (재배치 가능성) ===
|
||||
persistentvolumeclaim/postgres-data → kc-lab-2
|
||||
|
||||
=== 서비스 정상 확인 ===
|
||||
https://auth.hyeonworks.com/realms/master HTTP 200
|
||||
|
||||
=== VM 상태 ===
|
||||
--------------------------
|
||||
1 kc-lab-1 running
|
||||
2 kc-lab-2 running
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
=== 워커 노드(kc-lab-2) 전원 차단 — virsh destroy 는 종료 신호가 없다 ===
|
||||
차단 시각: 12:07:43
|
||||
Domain 'kc-lab-2' destroyed
|
||||
|
||||
|
||||
+15초 node=Ready | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 000
|
||||
+30초 node=Ready | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 000
|
||||
+45초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+60초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+75초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+90초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+105초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+120초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+135초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+150초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+165초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
+180초 node=NotReady | keycloak-0=Running postgres-7b474b88c8-2gf27=Running | 외부 HTTP 503
|
||||
@@ -0,0 +1,28 @@
|
||||
=== 파드 상태의 진실 — Running 인데 노드가 없다 ===
|
||||
a2-probe Running true kc-lab-2 <none>
|
||||
keycloak-0 Running true kc-lab-2 <none>
|
||||
keycloak-1 Running false kc-lab-1 <none>
|
||||
postgres-7b474b88c8-2gf27 Running true kc-lab-2 <none>
|
||||
|
||||
=== 재배치가 시도되었는가 ===
|
||||
10m Warning Unhealthy pod/keycloak-0 Readiness probe failed: Get "http://10.42.1.67:9000/health/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
|
||||
3m15s Warning NodeNotReady pod/postgres-7b474b88c8-2gf27 Node is not ready
|
||||
3m15s Warning NodeNotReady pod/keycloak-0 Node is not ready
|
||||
3m15s Warning NodeNotReady pod/a2-probe Node is not ready
|
||||
2m27s Warning Unhealthy pod/keycloak-1 Readiness probe failed: Get "http://10.42.0.35:9000/health/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
|
||||
2s Warning Unhealthy pod/keycloak-1 Readiness probe failed: HTTP probe failed with statuscode: 503
|
||||
|
||||
=== 노드 taint — 쿠버네티스가 붙인 것 ===
|
||||
node.kubernetes.io/unreachable=:NoSchedule
|
||||
node.kubernetes.io/unreachable=:NoExecute
|
||||
|
||||
=== Prometheus 가 본 것 (kc-lab-1 에 있어 살아남았다) ===
|
||||
up{job=keycloak pod=keycloak-1 } = 1
|
||||
up{job=keycloak pod=keycloak-0 } = 0
|
||||
up{job=kubelet pod=- } = 1
|
||||
up{job=kubelet pod=- } = 0
|
||||
up{job=node-exporter pod=kc-lab-1 } = 1
|
||||
up{job=node-exporter pod=kc-lab-2 } = 0
|
||||
up{job=prometheus pod=- } = 1
|
||||
|
||||
=== 진입점이 처음 40초간 000 이었던 이유 — nginx upstream ===
|
||||
@@ -0,0 +1,15 @@
|
||||
=== nginx 설정 위치 찾기 ===
|
||||
|
||||
=== NoExecute taint 의 tolerationSeconds — 언제 축출되는가 ===
|
||||
node.kubernetes.io/not-ready NoExecute tolerationSeconds=300
|
||||
node.kubernetes.io/unreachable NoExecute tolerationSeconds=300
|
||||
|
||||
=== 5분 축출 시점까지 관찰 ===
|
||||
+210초 a2-probe:Running keycloak-0:Running keycloak-1:Running postgres-7b474b88c8-2gf27:Running
|
||||
+240초 a2-probe:Running keycloak-0:Running keycloak-1:Running postgres-7b474b88c8-2gf27:Running
|
||||
+270초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
+300초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
+330초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
+360초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
+390초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
+420초 a2-probe:Terminating keycloak-0:Terminating keycloak-1:Running postgres-7b474b88c8-2gf27:Terminating postgres-7b474b88c8-9cmsv:Pending
|
||||
@@ -0,0 +1,18 @@
|
||||
=== 새 postgres 가 Pending 인 이유 ===
|
||||
Events:
|
||||
Type Reason Age From Message
|
||||
---- ------ ---- ---- -------
|
||||
Warning FailedScheduling 4m45s default-scheduler 0/2 nodes are available: 1 node(s) didn't match PersistentVolume's node affinity, 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling.
|
||||
|
||||
=== keycloak-0 대체 파드가 안 생기는 이유 (StatefulSet) ===
|
||||
keycloak 2 <none> 1
|
||||
keycloak-0 1/1 Terminating 0 30m
|
||||
keycloak-1 0/1 Running 0 143m
|
||||
|
||||
=== 복구 — 노드 재기동 ===
|
||||
재기동 시각: 12:16:31
|
||||
Domain 'kc-lab-2' started
|
||||
|
||||
+30초 node=Ready | Running 파드 3 개 | 외부 HTTP 503
|
||||
+60초 node=Ready | Running 파드 3 개 | 외부 HTTP 200
|
||||
→ 서비스 복귀
|
||||
@@ -0,0 +1,19 @@
|
||||
=== 복구 확인 ===
|
||||
keycloak-0 1/1 Running 0 68s
|
||||
keycloak-1 1/1 Running 0 144m
|
||||
postgres-7b474b88c8-9cmsv 1/1 Running 0 4m20s
|
||||
|
||||
=== kc-lab-1(k3s server)에 무엇이 있는가 — 이게 곧 영향 범위다 ===
|
||||
keycloak-lab keycloak-1
|
||||
kube-system coredns-54996dc9b4-8k8fj
|
||||
kube-system helm-install-traefik-crd-q29b5
|
||||
kube-system local-path-provisioner-77b9867795-g27z8
|
||||
kube-system metrics-server-6dc596dfb8-7xxq4
|
||||
kube-system svclb-traefik-5eb6a9a1-qwwk5
|
||||
kube-system traefik-5d6fcf895-wpfhr
|
||||
observability grafana-845b5678cf-b6gvc
|
||||
observability node-exporter-9qk9w
|
||||
observability prometheus-6774f94f7c-pzr2t
|
||||
|
||||
=== Traefik replica 수 (진입점의 단일 장애점인가) ===
|
||||
traefik 1 1
|
||||
@@ -0,0 +1,19 @@
|
||||
=== 컨트롤 플레인 노드(kc-lab-1) 전원 차단 ===
|
||||
차단 시각: 12:18:08
|
||||
Domain 'kc-lab-1' destroyed
|
||||
|
||||
|
||||
+20초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+40초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+60초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+80초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+100초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+120초 외부 auth=000 grafana=502 | kubectl: Unable to connect to the server: dial tcp
|
||||
+140초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
+160초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server: dial tcp
|
||||
|
||||
=== 살아 있는 노드에서 직접 확인 — 워크로드는 도는가 ===
|
||||
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
|
||||
e5f777900b762 60e153026e8f5 4 minutes ago Running keycloak 0 640d4dafaefb3 keycloak-0 keycloak-lab
|
||||
|
||||
6
|
||||
@@ -0,0 +1,13 @@
|
||||
=== 컨트롤 플레인 노드 복구 ===
|
||||
재기동: 12:23:39
|
||||
Domain 'kc-lab-1' started
|
||||
|
||||
+30초 외부=502 | kc-lab-1=Ready kc-lab-2=Ready
|
||||
+60초 외부=200 | kc-lab-1=Ready kc-lab-2=Ready
|
||||
→ 서비스 복귀 (총 60초)
|
||||
|
||||
=== 최종 상태 ===
|
||||
keycloak-0 1/1 Running 0 7m57s
|
||||
keycloak-1 1/1 Running 1 (<invalid> ago) 151m
|
||||
postgres-7b474b88c8-9cmsv 1/1 Running 0 11m
|
||||
(prometheus port-forward 재연결 필요)
|
||||
@@ -0,0 +1,23 @@
|
||||
# A-4 — 노드 전원 차단 증거
|
||||
|
||||
2026-09-04 12:07–12:24 KST
|
||||
해설: [`docs/experiment-a4-node-loss.md`](../../experiment-a4-node-loss.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-baseline.txt` | 차단 전 — 양쪽 Ready, **PVC 가 kc-lab-2 에 못박혀 있음**(재배치 불가의 원인), 외부 200 |
|
||||
| `02-worker-node-killed.txt` | `virsh destroy kc-lab-2` — **40초간 노드가 Ready 로 남아 있고** 외부는 이미 `000`. 이후 `503` |
|
||||
| `03-state-during-loss.txt` | **죽은 파드가 `ready=true`, 산 파드가 `ready=false`.** `up` 은 정확히 0. `unreachable` taint |
|
||||
| `04-eviction-timing.txt` | `tolerationSeconds=300` — **5분 뒤** 축출, 새 postgres 는 `Pending` |
|
||||
| `05-recovery.txt` | `FailedScheduling: didn't match PersistentVolume's node affinity`, StatefulSet `DESIRED=2 CURRENT=1`. 노드 복귀 후 **60초** |
|
||||
| `06-control-plane-inventory.txt` | kc-lab-1 에 있는 것 목록 — **Traefik `replicas=1`** |
|
||||
| `07-control-plane-loss.txt` | `virsh destroy kc-lab-1` — 외부 `000`, `kubectl` 불통. **그런데 `crictl ps` 로 보면 keycloak-0 은 Running** |
|
||||
| `08-control-plane-recovery.txt` | 60초 만에 복귀 |
|
||||
| `a4-up-dropped-per-node.png` | Grafana — `up` 이 노드별로 0 으로 떨어지는 구간. 12:18–12:23 은 **0 이 아니라 데이터 없음**(관측자가 같이 죽음) |
|
||||
|
||||
## 핵심 네 줄
|
||||
|
||||
1. **쿠버네티스는 40초 동안 노드가 살아 있다고 말한다.** 사용자는 이미 장애를 겪는 중이다.
|
||||
2. **죽은 파드의 상태는 화석이다.** `ready=true` 인 파드가 꺼진 기계 위에 있다.
|
||||
3. **StatefulSet 은 대체 파드를 만들지 않고, PVC 는 재배치를 막는다.** 사람이 개입해야 한다.
|
||||
4. **컨트롤 플레인 상실 ≠ 워크로드 상실.** 컨테이너는 계속 돌고, 들어갈 문만 사라진다.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 79 KiB |
@@ -0,0 +1,56 @@
|
||||
수집 시각: 2026-09-03 17:24:54 KST
|
||||
대상: Keycloak 26.7.0 × 2 + PostgreSQL 16, k3s 2노드
|
||||
|
||||
=== [1] 파드 배치 ===
|
||||
keycloak-0 1/1 10.42.1.18 kc-lab-2
|
||||
keycloak-1 1/1 10.42.0.16 kc-lab-1
|
||||
postgres-7b474b88c8-bw7b8 1/1 10.42.1.19 kc-lab-2
|
||||
|
||||
=== [2] 클러스터 뷰 로그 (Infinispan) ===
|
||||
-- keycloak-0 --
|
||||
2026-09-03 08:18:23,359 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938(v=16.0.12), keycloak-0-49501(v=16.0.12)]
|
||||
2026-09-03 08:18:23,433 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000079: Channel `ISPN` local address is `keycloak-0-49501`, physical addresses are `[10.42.1.18:7800]`
|
||||
-- keycloak-1 --
|
||||
2026-09-03 08:18:23,269 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938(v=16.0.12), keycloak-0-49501(v=16.0.12)]
|
||||
2026-09-03 08:18:23,282 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN100000: Node keycloak-0-49501 joined the cluster
|
||||
2026-09-03 08:18:23,286 INFO [org.infinispan.CLUSTER] (jgroups-5,keycloak-1-26938(v=16.0.12)) ISPN100000: Node keycloak-0-49501 joined the cluster
|
||||
|
||||
=== [3] JGROUPS_PING 테이블 구조 ===
|
||||
Table "public.jgroups_ping"
|
||||
Column | Type | Collation | Nullable | Default
|
||||
----------------+------------------------+-----------+----------+---------
|
||||
address | character varying(200) | | not null |
|
||||
name | character varying(200) | | |
|
||||
cluster_name | character varying(200) | | not null |
|
||||
ip | character varying(200) | | not null |
|
||||
coord | boolean | | |
|
||||
last_update | bigint | | |
|
||||
coordinated_by | character varying(200) | | |
|
||||
Indexes:
|
||||
"constraint_jgroups_ping" PRIMARY KEY, btree (address)
|
||||
|
||||
|
||||
=== [4] JGROUPS_PING 등록 내역 ===
|
||||
name | cluster_name | ip | coord
|
||||
------------------+--------------+-----------------+-------
|
||||
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
|
||||
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
|
||||
(2 rows)
|
||||
|
||||
|
||||
=== [5] 외부 접근 — OIDC discovery ===
|
||||
issuer https://auth.hyeonworks.com/realms/master
|
||||
authorization_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/auth
|
||||
token_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/token
|
||||
end_session_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/logout
|
||||
jwks_uri https://auth.hyeonworks.com/realms/master/protocol/openid-connect/certs
|
||||
|
||||
★ 전부 https. 첫 실험에서 확정한 KC_HOSTNAME + KC_PROXY_HEADERS 조합이 작동한다.
|
||||
|
||||
=== [6] 자원 사용 ===
|
||||
keycloak-0 8m 594Mi
|
||||
keycloak-1 9m 593Mi
|
||||
postgres-7b474b88c8-bw7b8 3m 67Mi
|
||||
--- 노드 ---
|
||||
kc-lab-1 2248Mi (65%)
|
||||
kc-lab-2 1447Mi (58%)
|
||||
@@ -0,0 +1,62 @@
|
||||
# 증거 — Keycloak 멀티노드 클러스터 형성
|
||||
|
||||
`docs/keycloak-multinode-cluster.md`의 근거 자료.
|
||||
**정상적으로 클러스터가 형성된 상태**에서 수집했으며, 이후 고장을 주입한
|
||||
뒤 이것과 대조한다.
|
||||
|
||||
수집 시각: 2026-09-03 17:24 KST
|
||||
|
||||
| 파일 | 내용 |
|
||||
|---|---|
|
||||
| `01-cluster-formed.txt` | 파드 배치·클러스터 뷰 로그·JGROUPS_PING·OIDC discovery·자원 |
|
||||
|
||||
## 이 상태에서 확인된 것
|
||||
|
||||
**클러스터 뷰가 멤버 2를 보고한다**
|
||||
|
||||
```
|
||||
ISPN000094: Received new cluster view for channel ISPN:
|
||||
[keycloak-1-26938|1] (2) [keycloak-1-26938, keycloak-0-49501]
|
||||
ISPN100000: Node keycloak-0-49501 joined the cluster
|
||||
ISPN000079: physical addresses are [10.42.1.18:7800]
|
||||
```
|
||||
|
||||
**디스커버리와 통신 경로가 한 테이블에 다 보인다**
|
||||
|
||||
```
|
||||
name | cluster_name | ip | coord
|
||||
------------------+--------------+-----------------+-------
|
||||
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
|
||||
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
|
||||
```
|
||||
|
||||
`name`/`cluster_name`은 **DB 디스커버리**의 결과이고, `ip`의 `:7800`은
|
||||
**실제 통신 경로**다. 7800을 막으면 이 표는 그대로 채워지면서 클러스터 뷰만
|
||||
깨질 것으로 예상한다 — 다음 실험의 가설이다.
|
||||
|
||||
`coord = t` 인 `keycloak-1`이 코디네이터다.
|
||||
|
||||
**배치** — 서로 다른 노드에 하나씩. PostgreSQL은 `kc-lab-2`에 있으므로
|
||||
**그 노드를 죽이면 Keycloak 하나와 DB가 동시에 사라진다.**
|
||||
|
||||
```
|
||||
keycloak-0 10.42.1.18 kc-lab-2
|
||||
keycloak-1 10.42.0.16 kc-lab-1
|
||||
postgres 10.42.1.19 kc-lab-2
|
||||
```
|
||||
|
||||
**issuer가 https로 발급된다** — 첫 실험(2홉 헤더 계약)의 결론이 적용된 결과다.
|
||||
|
||||
## 재수집
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab get pods -o wide
|
||||
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
|
||||
|
||||
PG=$(kubectl -n keycloak-lab get pod -l app=postgres -o name | head -1)
|
||||
kubectl -n keycloak-lab exec "$PG" -- \
|
||||
psql -U keycloak -d keycloak -c "SELECT name, cluster_name, ip, coord FROM jgroups_ping ORDER BY name;"
|
||||
|
||||
curl -s https://auth.hyeonworks.com/realms/master/.well-known/openid-configuration | python3 -m json.tool
|
||||
kubectl -n keycloak-lab top pods
|
||||
```
|
||||
@@ -0,0 +1,52 @@
|
||||
===================================================================
|
||||
실험 0 — 한 노드에서 만든 세션이 다른 노드에서 쓰이는가
|
||||
===================================================================
|
||||
|
||||
### 사전 확인: 클러스터가 2 멤버로 형성되었는가
|
||||
2026-09-04 00:52:09,294 INFO [org.infinispan.CLUSTER] (executor-thread-1) ISPN000094: Received new cluster view for channel ISPN: [keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749(v=16.0.12), keycloak-0-30843(v=16.0.12)]
|
||||
name | ip | coord
|
||||
------------------+-----------------+-------
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t
|
||||
keycloak-0-30843 | 10.42.1.43:7800 | f
|
||||
(2 rows)
|
||||
|
||||
|
||||
수집 시각: 2026-09-04 09:54:29 KST
|
||||
|
||||
=== 대상 ===
|
||||
keycloak-0 10.42.1.43 kc-lab-2
|
||||
keycloak-1 10.42.0.35 kc-lab-1
|
||||
|
||||
=== [0] 실험 전 DB 세션 ===
|
||||
|
||||
=== [1] keycloak-0 에서 로그인 ===
|
||||
sid jiv3rVZi1VeaO07oVJkL_MYW
|
||||
sub None
|
||||
iss https://auth.hyeonworks.com/realms/master
|
||||
access 수명 60초
|
||||
refresh 수명 1800초 typ=Refresh
|
||||
refresh jti 7669cc49-4778-851f-3c49-65f76964ae8e
|
||||
|
||||
=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===
|
||||
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
|
||||
ipAddress=10.42.1.44 start=1788483164000 lastAccess=1788483164000
|
||||
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
|
||||
ipAddress=10.42.1.44 start=1788483164000 lastAccess=1788483164000
|
||||
|
||||
=== [5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용 ===
|
||||
HTTP 200 ← 기대대로
|
||||
새 토큰의 sid → 동일 ✔
|
||||
|
||||
=== [6] keycloak-1 을 통해 로그아웃 ===
|
||||
http_code=204
|
||||
|
||||
=== [7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파) ===
|
||||
HTTP 400 ← 기대대로
|
||||
error invalid_grant
|
||||
error_description Session not active
|
||||
|
||||
=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ===
|
||||
대상 sid: jiv3rVZi1VeaO07oVJkL_MYW
|
||||
행 없음 — 로그아웃으로 삭제되었다
|
||||
|
||||
전체 세션 수: 1
|
||||
@@ -0,0 +1,35 @@
|
||||
===================================================================
|
||||
실험 0b — Infinispan 이 복제한 것인가, DB 를 같이 본 것인가
|
||||
===================================================================
|
||||
|
||||
수집 시각: 2026-09-04 09:54:41 KST
|
||||
|
||||
=== 로그인은 keycloak-0 에만 보냈다 ===
|
||||
로그인 응답: http_code=200
|
||||
|
||||
=== keycloak-0 (로그인을 받은 노드) ===
|
||||
계수기 캐시 전 후 증가
|
||||
rpc.replication_count clientSessions 1 1 +0
|
||||
rpc.replication_count sessions 1 1 +0
|
||||
approximate_entries_unique clientSessions 1 2 +1 ←
|
||||
approximate_entries_unique sessions 1 2 +1 ←
|
||||
hits clientSessions 2 2 +0
|
||||
hits sessions 2 2 +0
|
||||
misses clientSessions 2 3 +1 ←
|
||||
misses sessions 3 4 +1 ←
|
||||
stores clientSessions 2 3 +1 ←
|
||||
stores sessions 2 3 +1 ←
|
||||
|
||||
=== keycloak-1 (아무 요청도 받지 않은 노드) ===
|
||||
계수기 캐시 전 후 증가
|
||||
rpc.replication_count clientSessions 7 7 +0
|
||||
rpc.replication_count sessions 7 7 +0
|
||||
approximate_entries_unique clientSessions 0 0 +0
|
||||
approximate_entries_unique sessions 0 0 +0
|
||||
hits clientSessions 4 4 +0
|
||||
hits sessions 4 4 +0
|
||||
misses clientSessions 0 0 +0
|
||||
misses sessions 0 0 +0
|
||||
stores clientSessions 1 1 +0
|
||||
stores sessions 1 1 +0
|
||||
|
||||
@@ -0,0 +1,15 @@
|
||||
===================================================================
|
||||
실험 0c — 세션 엔트리는 어느 노드에 있는가 (로컬 캐시인가 분산인가)
|
||||
===================================================================
|
||||
|
||||
수집 시각: 2026-09-04 09:54:54 KST
|
||||
keycloak-0 = 10.42.1.43 (kc-lab-2)
|
||||
keycloak-1 = 10.42.0.35 (kc-lab-1)
|
||||
|
||||
단계 k0 entries k1 entries
|
||||
시작 2.0 0.0
|
||||
keycloak-1 에 로그인 5회 2.0 5.0
|
||||
keycloak-0 에 로그인 5회 7.0 5.0
|
||||
|
||||
=== 대조: PostgreSQL 에는 몇 건인가 ===
|
||||
online 세션 12
|
||||
@@ -0,0 +1,56 @@
|
||||
===================================================================
|
||||
실험 0d — 반대편 노드가 정말 DB 에서 읽는가 (SQL 을 직접 잡는다)
|
||||
===================================================================
|
||||
|
||||
수집 시각: 2026-09-04 10:14:17 KST
|
||||
keycloak-0 = 10.42.1.43 (세션을 만드는 노드)
|
||||
keycloak-1 = 10.42.0.35 (읽기만 하는 노드)
|
||||
|
||||
=== PostgreSQL 문장 로깅을 켠다 ===
|
||||
log_statement = all
|
||||
log_line_prefix = %m [%p] %h
|
||||
|
||||
=== 요청 ===
|
||||
SID=jSt9GEPVQLJsO-1CeJjVgltg
|
||||
K1_ENTRIES_BEFORE=5.0
|
||||
REFRESH_ON_K1=200
|
||||
K1_ENTRIES_AFTER=5.0
|
||||
|
||||
=== PostgreSQL 문장 로깅을 끈다 ===
|
||||
log_statement = none
|
||||
|
||||
=== keycloak-1 이 실제로 보낸 SQL 문장 ===
|
||||
(파라미터가 $1 로 묶여 있어, sid 는 바로 아래 DETAIL 줄에 있다)
|
||||
|
||||
select puse1_0.OFFLINE_FLAG,puse1_0.USER_SESSION_ID,puse1_0.BROKER_SESSION_ID,puse1_0.CREATED_ON,puse1_0.DATA,puse1_0.LAST_SESSION_REFRESH,puse1_0.REALM_ID,puse1_0.REMEMBER_ME,puse1_0.USER_ID,puse1_0.VERSION from OFFLINE_USER_SESSION puse1_0 where (puse1_0.OFFLINE_FLAG,puse1_0.USER_SESSION_ID) in (($1,$2))
|
||||
select puse1_0.VERSION from OFFLINE_USER_SESSION puse1_0 where puse1_0.USER_SESSION_ID=$1 and puse1_0.OFFLINE_FLAG=$2 for no key update of puse1_0 skip locked
|
||||
select pcse1_0.CLIENT_ID,pcse1_0.CLIENT_STORAGE_PROVIDER,pcse1_0.EXTERNAL_CLIENT_ID,pcse1_0.OFFLINE_FLAG,pcse1_0.USER_SESSION_ID,pcse1_0.DATA,pcse1_0.REALM_ID,pcse1_0.TIMESTAMP,pcse1_0.VERSION from OFFLINE_CLIENT_SESSION pcse1_0 where (pcse1_0.CLIENT_ID,pcse1_0.CLIENT_STORAGE_PROVIDER,pcse1_0.EXTERNAL_CLIENT_ID,pcse1_0.OFFLINE_FLAG,pcse1_0.USER_SESSION_ID) in (($1,$2,$3,$4,$5))
|
||||
select pcse1_0.VERSION from OFFLINE_CLIENT_SESSION pcse1_0 where pcse1_0.USER_SESSION_ID=$1 and pcse1_0.OFFLINE_FLAG=$2 and pcse1_0.CLIENT_ID=$3 and pcse1_0.EXTERNAL_CLIENT_ID=$4 and pcse1_0.CLIENT_STORAGE_PROVIDER=$5 for no key update of pcse1_0 skip locked
|
||||
update OFFLINE_CLIENT_SESSION set TIMESTAMP=$1,VERSION=$2 where CLIENT_ID=$3 and CLIENT_STORAGE_PROVIDER=$4 and EXTERNAL_CLIENT_ID=$5 and OFFLINE_FLAG=$6 and USER_SESSION_ID=$7 and VERSION=$8
|
||||
update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1,VERSION=$2 where OFFLINE_FLAG=$3 and USER_SESSION_ID=$4 and VERSION=$5
|
||||
SET LOCAL synchronous_commit TO OFF
|
||||
COMMIT
|
||||
DELETE from JGROUPS_PING WHERE address=$1
|
||||
INSERT INTO JGROUPS_PING (address, name, cluster_name, ip, coord, last_update, coordinated_by) values ($1, $2, $3, $4, $5, $6, $7)
|
||||
COMMIT
|
||||
DELETE from JGROUPS_PING WHERE address=$1
|
||||
|
||||
=== 그 sid 를 언급한 SQL — 누가 보냈는가 ===
|
||||
찾는 sid: jSt9GEPVQLJsO-1CeJjVgltg
|
||||
|
||||
2026-09-04 01:12:32.851 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:32.852 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:32.860 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:32.862 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:32.863 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = NULL, $2 = '1788484352', $3 = '{"ipAddress":"10.42.1.50","authMethod":"openid-connect","rememberMe":false,"started":0,"notes":{"KC_DEVICE_NOTE":"
|
||||
2026-09-04 01:12:32.864 UTC [81407] [keycloak-0] DETAIL: parameters: $1 = '{"authMethod":"openid-connect","notes":{"clientId":"131a9912-b578-4b9c-b16a-97518704077e","userSessionStartedAt":"1788484352","iss":"https://aut
|
||||
2026-09-04 01:12:34.934 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '0', $2 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:34.936 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = 'jSt9GEPVQLJsO-1CeJjVgltg', $2 = '0'
|
||||
2026-09-04 01:12:34.937 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '131a9912-b578-4b9c-b16a-97518704077e', $2 = 'local', $3 = 'local', $4 = '0', $5 = 'jSt9GEPVQLJsO-1CeJjVgltg'
|
||||
2026-09-04 01:12:34.938 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = 'jSt9GEPVQLJsO-1CeJjVgltg', $2 = '0', $3 = '131a9912-b578-4b9c-b16a-97518704077e', $4 = 'local', $5 = 'local'
|
||||
2026-09-04 01:12:34.944 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '1788484354', $2 = '1', $3 = '131a9912-b578-4b9c-b16a-97518704077e', $4 = 'local', $5 = 'local', $6 = '0', $7 = 'jSt9GEPVQLJsO-1CeJjVgltg', $8 =
|
||||
2026-09-04 01:12:34.946 UTC [81376] [keycloak-1] DETAIL: parameters: $1 = '1788484354', $2 = '1', $3 = '0', $4 = 'jSt9GEPVQLJsO-1CeJjVgltg', $5 = '0'
|
||||
|
||||
=== 요약: 파드별 질의 건수 ===
|
||||
6 [keycloak-1]
|
||||
6 [keycloak-0]
|
||||
@@ -0,0 +1,26 @@
|
||||
# 실험 0 — 세션 복제 증거
|
||||
|
||||
수집: 2026-09-04 09:54 KST · Keycloak 26 / Infinispan 16.0.12 / PostgreSQL 16
|
||||
해설: [`docs/experiment-00-session-replication.md`](../../experiment-00-session-replication.md)
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `01-cross-node-session.txt` | 클러스터 2멤버 확인 → keycloak-0 로그인 → 같은 sid 가 양쪽에서 보임 → **keycloak-1 이 refresh 성공(200)** → keycloak-1 로그아웃 → **keycloak-0 갱신 실패(400)** → DB 행 삭제 확인 |
|
||||
| `02-cache-delta.txt` | 로그인 하나를 사이에 둔 양쪽 노드의 캐시 계수기. **keycloak-1 은 전부 +0** |
|
||||
| `03-cache-ownership.txt` | 로그인을 반대편에 몰아준 결과. **요청을 받은 노드에서만 엔트리가 는다.** 캐시 합 7+5 = DB 12 |
|
||||
| `session-cache-entries-per-pod.png` | 위 사실의 시계열. 파란 선(keycloak-1)이 0에 붙어 있는 동안 초록 선(keycloak-0)만 14까지 오른다 |
|
||||
| `keycloak-admin-sessions.png` | 관리 콘솔의 Sessions 화면. 브라우저는 nginx→Traefik 을 거쳐 두 파드 중 하나에 닿지만 **어느 파드가 만든 세션이든 전부 보인다** |
|
||||
|
||||
## 핵심 한 줄
|
||||
|
||||
클러스터는 형성되지만 **세션 엔트리는 노드를 건너가지 않는다.**
|
||||
두 노드가 같은 답을 하는 이유는 Infinispan 복제가 아니라 **같은 PostgreSQL** 이다.
|
||||
|
||||
| 파일 | 무엇을 보여주는가 |
|
||||
|---|---|
|
||||
| `04-read-path-sql.txt` | PostgreSQL 문장 로깅으로 잡은 **keycloak-1 이 실제로 날린 SQL**. `SELECT ... FROM OFFLINE_USER_SESSION` 로 남의 세션을 읽고 `UPDATE ... where VERSION=$5` 로 쓴다. 같은 트랜잭션에 `SET LOCAL synchronous_commit TO OFF` 가 들어 있다 |
|
||||
|
||||
## 추론이 관측이 된 지점
|
||||
|
||||
0b·0c 는 "keycloak-1 메모리에 없는데 쓸 수 있으니 DB 에서 읽었을 것"이라는
|
||||
**추론**이었다. 0d 에서 그 SQL 을 파드 IP 와 함께 직접 잡았다.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 116 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 76 KiB |
@@ -0,0 +1,636 @@
|
||||
# 실험 0 — 한 노드에서 만든 세션이 다른 노드에서 쓰이는가
|
||||
|
||||
로드맵 A-0. 이후 모든 장애 실험의 기준선이다.
|
||||
|
||||
> **맥락이 안 잡히면 먼저 읽을 것** —
|
||||
> [`docs/session-lab-prerequisites.md`](session-lab-prerequisites.md).
|
||||
> 왜 세션이 문제가 되는지, Keycloak이 세션을 어디에 두는지, 그래서 이 실험이
|
||||
> 무엇을 가르려는 것인지를 바닥부터 세워둔 문서다.
|
||||
|
||||
- 실행 스크립트 — [`deploy/lab/scripts/experiment-session-replication.sh`](../deploy/lab/scripts/experiment-session-replication.sh),
|
||||
[`experiment-cache-replication-delta.sh`](../deploy/lab/scripts/experiment-cache-replication-delta.sh),
|
||||
[`experiment-cache-ownership.sh`](../deploy/lab/scripts/experiment-cache-ownership.sh),
|
||||
[`experiment-session-read-path.sh`](../deploy/lab/scripts/experiment-session-read-path.sh)
|
||||
- 증거 — [`docs/evidence/session-replication/`](evidence/session-replication/)
|
||||
- 수집 시각 — 2026-09-04 09:54 KST, Keycloak 26 / Infinispan 16.0.12 / PostgreSQL 16
|
||||
|
||||
---
|
||||
|
||||
## 0. 결론부터
|
||||
|
||||
| 물음 | 답 |
|
||||
|---|---|
|
||||
| 한 노드에서 만든 세션을 다른 노드가 쓸 수 있는가 | **그렇다** |
|
||||
| 로그아웃이 반대 방향으로 전파되는가 | **그렇다** |
|
||||
| **그 공유는 Infinispan 복제 덕분인가** | **아니다** |
|
||||
| 그럼 무엇이 공유하는가 | **PostgreSQL** — 반대편 노드가 날린 SQL을 직접 잡았다 |
|
||||
|
||||
**클러스터가 형성됐다는 것과 세션이 복제된다는 것은 다른 얘기였다.**
|
||||
로그에는 `(2) [keycloak-0, keycloak-1]`이 찍히고 `JGROUPS_PING`에도 둘 다
|
||||
등록되어 있지만, **세션 엔트리는 노드 사이를 건너가지 않는다.**
|
||||
|
||||
각 노드는 **자기가 처리한 로그인만** 캐시한다. 두 노드가 같은 답을 내놓는
|
||||
이유는 복제가 아니라 **같은 데이터베이스를 보기 때문**이다.
|
||||
|
||||
---
|
||||
|
||||
## 1. 왜 이 실험이 첫 번째인가
|
||||
|
||||
앞선 작업에서 Keycloak 2노드 클러스터를 세우고 `ISPN000094`로 멤버 2개를
|
||||
확인했다. 거기서 멈추면 **"클러스터가 떴다"까지만 아는 것**이고, 그 위에서
|
||||
장애를 주입해봐야 무엇이 무엇 때문에 깨졌는지 해석할 수 없다.
|
||||
|
||||
기준선이 없으면 이런 잘못된 추론을 하게 된다.
|
||||
|
||||
> 7800을 막았더니 세션이 깨졌다 → 역시 세션은 7800으로 복제되는구나
|
||||
|
||||
실제로는 7800으로 세션이 오가지 않는다는 것을 **먼저** 알아야, 7800을 막았을
|
||||
때 깨지는 것이 무엇인지 정확히 말할 수 있다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 실험 설계에서 배운 것 세 가지
|
||||
|
||||
측정값보다 **어떻게 측정할지**에서 더 많이 틀렸다. 세 번 고쳤다.
|
||||
|
||||
### 2-1. 대조군 없는 측정은 해석할 수 없다
|
||||
|
||||
첫 판본은 이렇게 보고했다.
|
||||
|
||||
```
|
||||
=== [4] keycloak-0 이 발급한 토큰을 keycloak-1 이 받는가 ===
|
||||
http_code=403
|
||||
```
|
||||
|
||||
**403을 "복제 실패"로 읽을 뻔했다.** 발급 노드에도 같은 요청을 보내보니
|
||||
|
||||
```
|
||||
--- userinfo, scope 없음 ---
|
||||
k0(발급노드) 403
|
||||
k1(반대편) 403
|
||||
--- 403 본문 ---
|
||||
WWW-Authenticate: Bearer realm="master", error="insufficient_scope",
|
||||
error_description="Missing openid scope"
|
||||
```
|
||||
|
||||
**양쪽 다 403이었다.** 원인은 복제가 아니라 요청에 `openid` scope가 없다는
|
||||
것이었다. 오히려 **두 노드가 똑같이 답했다는 사실 자체가 일치의 증거**였다.
|
||||
|
||||
> **원칙** — 반대편 노드의 응답은 발급 노드의 응답과 나란히 놓기 전까지
|
||||
> 아무 의미가 없다. 시험군만 재는 측정은 측정이 아니다.
|
||||
|
||||
### 2-2. 개수가 아니라 식별자로 추적한다
|
||||
|
||||
`client-session-stats`가 `active=2`를 돌려줬다. 그런데 스크립트 자체가
|
||||
로그인을 두 번 하고(시험용 + 관리 API 호출용) 있었다. **개수는 실험 도구가
|
||||
만든 잡음에 그대로 오염된다.**
|
||||
|
||||
바꾼 방식: 토큰의 `sid`를 뽑아, 각 노드의 세션 목록에 **그 sid가 있는지**를
|
||||
본다. 개수가 몇이든 상관없다.
|
||||
|
||||
```
|
||||
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
|
||||
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
|
||||
```
|
||||
|
||||
### 2-3. 세션 저장소를 실제로 건드리는 탐침을 골라야 한다
|
||||
|
||||
| 탐침 | 하는 일 | 적합한가 |
|
||||
|---|---|---|
|
||||
| `userinfo` | 서명 검증 + scope 확인 | **아니다.** 세션을 몰라도 통과할 수 있다 |
|
||||
| **`refresh_token` 그랜트** | 세션을 찾고, 살아있는지 보고, 갱신 시각을 쓴다 | **그렇다** |
|
||||
|
||||
refresh는 **읽고 쓴다.** 그래서 "저 노드가 이 세션을 정말로 아는가"에 답한다.
|
||||
|
||||
여기에 더해 refresh token은 **회전(rotation)** 된다 — 한 번 쓰면 옛 것이
|
||||
무효가 된다. 따라서 **반대편 노드에 먼저 써야** 한다. 발급 노드에 먼저 쓰면
|
||||
시험군에 쓸 토큰이 사라진다. 대조군과 시험군의 순서가 강제된다.
|
||||
|
||||
---
|
||||
|
||||
## 3. 실험 0 — 교차 노드 세션 사용
|
||||
|
||||
### 실행
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "delete from offline_user_session"
|
||||
kubectl -n keycloak-lab rollout restart statefulset/keycloak # 캐시를 비운다
|
||||
./deploy/lab/scripts/experiment-session-replication.sh
|
||||
```
|
||||
|
||||
nginx나 Traefik을 거치지 않고 **파드 IP로 직접** 말을 건다. 로드밸런서를
|
||||
거치면 어느 노드가 처리했는지가 감춰지는데, 그게 바로 이 실험의 질문이다.
|
||||
|
||||
### 결과 — [`01-cross-node-session.txt`](evidence/session-replication/01-cross-node-session.txt)
|
||||
|
||||
```
|
||||
### 사전 확인: 클러스터가 2 멤버로 형성되었는가
|
||||
ISPN000094: Received new cluster view for channel ISPN:
|
||||
[keycloak-1-48749(v=16.0.12)|5] (2) [keycloak-1-48749, keycloak-0-30843]
|
||||
|
||||
name | ip | coord
|
||||
------------------+-----------------+-------
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t
|
||||
keycloak-0-30843 | 10.42.1.43:7800 | f
|
||||
|
||||
=== 대상 ===
|
||||
keycloak-0 10.42.1.43 kc-lab-2
|
||||
keycloak-1 10.42.0.35 kc-lab-1
|
||||
|
||||
=== [1] keycloak-0 에서 로그인 ===
|
||||
sid jiv3rVZi1VeaO07oVJkL_MYW
|
||||
iss https://auth.hyeonworks.com/realms/master
|
||||
access 수명 60초
|
||||
refresh 수명 1800초 typ=Refresh
|
||||
|
||||
=== [3] 같은 sid 가 두 노드 모두에서 보이는가 ===
|
||||
keycloak-0 (발급 노드) 세션 2개 중 대상 sid → 보임 ✔
|
||||
keycloak-1 (반대편) 세션 2개 중 대상 sid → 보임 ✔
|
||||
|
||||
=== [5] keycloak-0 이 발급한 refresh token 을 keycloak-1 에 사용 ===
|
||||
HTTP 200 ← 기대대로
|
||||
새 토큰의 sid → 동일 ✔
|
||||
|
||||
=== [6] keycloak-1 을 통해 로그아웃 ===
|
||||
http_code=204
|
||||
|
||||
=== [7] 로그아웃 후 keycloak-0 에서 갱신 시도 (무효화 전파) ===
|
||||
HTTP 400 ← 기대대로
|
||||
error invalid_grant
|
||||
error_description Session not active
|
||||
|
||||
=== [8] PostgreSQL 에서 그 sid 를 직접 확인 ===
|
||||
행 없음 — 로그아웃으로 삭제되었다
|
||||
```
|
||||
|
||||
**네 가지가 모두 기대대로다.**
|
||||
|
||||
| | 확인된 것 |
|
||||
|---|---|
|
||||
| 조회 | 같은 sid가 양쪽에서 보인다 |
|
||||
| **쓰기** | keycloak-0의 refresh token을 keycloak-1이 받아 갱신했고, **sid가 유지된다** |
|
||||
| **역방향 무효화** | keycloak-1의 로그아웃이 keycloak-0의 갱신을 막았다 |
|
||||
| 영속 | 로그아웃과 함께 DB 행이 사라졌다 |
|
||||
|
||||
**`sid`는 JWT 안에만 있는 값이 아니다.** PostgreSQL의
|
||||
`OFFLINE_USER_SESSION.user_session_id` 컬럼에 **문자 그대로** 들어 있다.
|
||||
|
||||
---
|
||||
|
||||
## 4. 실험 0b — 복제인가, 같은 DB를 본 것인가
|
||||
|
||||
실험 0은 "두 노드가 같은 답을 한다"까지만 증명한다. **그것으로는 Infinispan이
|
||||
복제했다고 말할 수 없다.** `persistent-user-sessions`(Keycloak 26 기본값)에서는
|
||||
세션이 PostgreSQL에 기록되므로, **캐시를 아예 꺼도 두 노드는 같은 답을 한다.**
|
||||
|
||||
가르는 방법: 로그인 한 번을 사이에 두고 **양쪽 노드의 캐시 계수기**를 잰다.
|
||||
|
||||
### 결과 — [`02-cache-delta.txt`](evidence/session-replication/02-cache-delta.txt)
|
||||
|
||||
```
|
||||
=== 로그인은 keycloak-0 에만 보냈다 ===
|
||||
로그인 응답: http_code=200
|
||||
|
||||
=== keycloak-0 (로그인을 받은 노드) ===
|
||||
계수기 캐시 전 후 증가
|
||||
approximate_entries_unique sessions 1 2 +1 ←
|
||||
stores sessions 2 3 +1 ←
|
||||
misses sessions 3 4 +1 ←
|
||||
rpc.replication_count sessions 1 1 +0
|
||||
|
||||
=== keycloak-1 (아무 요청도 받지 않은 노드) ===
|
||||
approximate_entries_unique sessions 0 0 +0
|
||||
stores sessions 1 1 +0
|
||||
hits sessions 4 4 +0
|
||||
rpc.replication_count sessions 7 7 +0
|
||||
```
|
||||
|
||||
**keycloak-1의 계수기가 하나도 움직이지 않았다.** 엔트리도 0, 저장도 0.
|
||||
|
||||
그리고 keycloak-1의 `sessions` 캐시 엔트리는 **처음부터 끝까지 0**이다.
|
||||
keycloak-0이 세션을 9개 들고 있는 동안에도 0이었다.
|
||||
|
||||
---
|
||||
|
||||
## 5. 실험 0c — 엔트리는 어느 노드에 있는가
|
||||
|
||||
0b의 결과에는 두 가지 설명이 가능하다.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| (a) **분산 캐시 + owners=1** | 일관 해싱으로 흩어지는데 이번 건이 우연히 keycloak-0에 떨어졌다 |
|
||||
| (b) **로컬 캐시** | 각 노드는 자기가 처리한 것만 캐시한다 |
|
||||
|
||||
**반대편 노드에 로그인을 몰아주면 갈린다.** (a)라면 어느 쪽에 요청하든 엔트리는
|
||||
양쪽에 흩어진다. (b)라면 **요청을 받은 노드에서만** 는다.
|
||||
|
||||
### 결과 — [`03-cache-ownership.txt`](evidence/session-replication/03-cache-ownership.txt)
|
||||
|
||||
```
|
||||
단계 k0 entries k1 entries
|
||||
시작 2.0 0.0
|
||||
keycloak-1 에 로그인 5회 2.0 5.0 ← k0 그대로, k1 만 +5
|
||||
keycloak-0 에 로그인 5회 7.0 5.0 ← k0 만 +5, k1 그대로
|
||||
|
||||
=== 대조: PostgreSQL 에는 몇 건인가 ===
|
||||
online 세션 12 ← 7 + 5 = 12, 정확히 일치
|
||||
```
|
||||
|
||||
**(b)다.** 그리고 **7 + 5 = 12**로 DB 총계와 정확히 맞는다 — 모든 세션이 DB에
|
||||
있고, 각각은 **자기를 만든 노드 한 곳에만** 캐시되어 있다.
|
||||
|
||||
### 그래프로 본 같은 사실
|
||||
|
||||

|
||||
|
||||
`vendor_statistics_approximate_entries_unique{cache="sessions"}` — Grafana Explore.
|
||||
|
||||
**파란 선(keycloak-1)이 0에 붙어 있는 동안 초록 선(keycloak-0)만 14까지
|
||||
올라간다.** 파란 선은 09:50, 즉 **keycloak-1에 직접 로그인을 보낸 순간에만**
|
||||
5로 뛴다. 중간의 절벽은 캐시를 비우려고 파드를 재시작한 지점이다.
|
||||
|
||||
> 캐시 설정은 파일에서 읽을 수 없다. 파드의 `/opt/keycloak/conf/cache-ispn.xml`은
|
||||
> `<cache-container name="keycloak"><transport/></cache-container>` 뿐이고,
|
||||
> Keycloak 26은 캐시를 **코드에서** 만든다. 그래서 위 결론은 설정을 읽어서가
|
||||
> 아니라 **동작을 측정해서** 얻었다.
|
||||
|
||||
---
|
||||
|
||||
## 6. 실험 0d — 반대편 노드가 정말 DB에서 읽는가
|
||||
|
||||
0b·0c까지는 **추론**이었다. "keycloak-1의 메모리에 없는데 쓸 수 있으니 DB에서
|
||||
읽었을 것이다" — 그럴듯하지만 **SQL을 본 적은 없다.**
|
||||
|
||||
PostgreSQL의 문장 로깅을 몇 초만 켜고, keycloak-0에서 만든 세션에 대해
|
||||
**keycloak-1에 refresh를 딱 한 번** 보낸 뒤 로그를 뒤졌다.
|
||||
|
||||
```bash
|
||||
alter system set log_statement='all';
|
||||
alter system set log_line_prefix='%m [%p] %h '; -- %h 로 파드 IP 를 남긴다
|
||||
select pg_reload_conf();
|
||||
```
|
||||
|
||||
### 잡힌 트랜잭션 — [`04-read-path-sql.txt`](evidence/session-replication/04-read-path-sql.txt)
|
||||
|
||||
```
|
||||
01:12:34.934 pid=81376 | BEGIN
|
||||
01:12:34.934 pid=81376 | select ... from OFFLINE_USER_SESSION where (OFFLINE_FLAG,USER_SESSION_ID) in (($1,$2))
|
||||
01:12:34.936 pid=81376 | select VERSION from OFFLINE_USER_SESSION ... for no key update skip locked
|
||||
01:12:34.937 pid=81376 | select ... from OFFLINE_CLIENT_SESSION where (...) in ((...))
|
||||
01:12:34.938 pid=81376 | select VERSION from OFFLINE_CLIENT_SESSION ... for no key update skip locked
|
||||
01:12:34.944 pid=81376 | update OFFLINE_CLIENT_SESSION set TIMESTAMP=$1,VERSION=$2 where ... and VERSION=$8
|
||||
01:12:34.946 pid=81376 | update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1,VERSION=$2 where ... and VERSION=$5
|
||||
01:12:34.946 pid=81376 | SET LOCAL synchronous_commit TO OFF
|
||||
01:12:34.947 pid=81376 | COMMIT
|
||||
```
|
||||
|
||||
이 연결의 클라이언트 IP는 `10.42.0.35` — **keycloak-1의 파드 IP**다.
|
||||
sid 하나에 대해 keycloak-0이 6건(로그인), keycloak-1이 6건(갱신)을 날렸다.
|
||||
|
||||
```
|
||||
=== 요약: 파드별 질의 건수 ===
|
||||
6 [keycloak-1]
|
||||
6 [keycloak-0]
|
||||
```
|
||||
|
||||
**추론이 관측이 되었다.** keycloak-1은 세션을 DB에서 읽고, DB에 쓴다.
|
||||
|
||||
### 여기서 딸려 나온 것 세 가지
|
||||
|
||||
이 13밀리초짜리 트랜잭션 하나에 **원래 질문들의 답이 절반쯤 들어 있다.**
|
||||
|
||||
#### (1) 낙관적 락 — `VERSION` 컬럼
|
||||
|
||||
```sql
|
||||
update OFFLINE_USER_SESSION
|
||||
set LAST_SESSION_REFRESH=$1, VERSION=$2
|
||||
where OFFLINE_FLAG=$3 and USER_SESSION_ID=$4 and VERSION=$5
|
||||
─────────────
|
||||
읽을 때의 버전과 같을 때만 쓴다
|
||||
```
|
||||
|
||||
읽은 뒤 다른 노드가 먼저 고쳤다면 `VERSION`이 달라져 **`UPDATE`가 0행을
|
||||
갱신하고 실패한다.** 잠금을 오래 잡지 않고 충돌을 사후에 검출하는 방식이다.
|
||||
|
||||
**리프레시 토큰 동시 갱신 경쟁(로드맵 B-5)이 여기서 갈린다.** 두 요청이
|
||||
같은 세션을 동시에 갱신하면 하나는 이 검사에서 진다.
|
||||
|
||||
#### (2) `FOR NO KEY UPDATE ... SKIP LOCKED`
|
||||
|
||||
```sql
|
||||
select VERSION from OFFLINE_USER_SESSION
|
||||
where USER_SESSION_ID=$1 and OFFLINE_FLAG=$2
|
||||
for no key update of puse1_0 skip locked
|
||||
────────────── ────────────
|
||||
키가 아닌 컬럼만 잠근다 잠긴 행은 건너뛴다 (기다리지 않는다)
|
||||
```
|
||||
|
||||
| 절 | 뜻 |
|
||||
|---|---|
|
||||
| `FOR NO KEY UPDATE` | 행을 잠그되 **외래키 참조는 막지 않는다.** `FOR UPDATE`보다 약해 경합이 준다 |
|
||||
| **`SKIP LOCKED`** | 이미 잠긴 행을 **기다리지 않고 건너뛴다** |
|
||||
|
||||
`SKIP LOCKED`가 핵심이다. 같은 세션에 동시 요청이 몰려도 **줄을 서지 않는다.**
|
||||
대기 대신 낙관적 락 실패로 처리한다 — 처리량을 위해 **지연 대신 재시도**를
|
||||
고른 설계다.
|
||||
|
||||
#### (3) `SET LOCAL synchronous_commit TO OFF` — 내구성을 일부 포기한다
|
||||
|
||||
**같은 트랜잭션 안에서**, `COMMIT` 직전에 나온다. pid로 경계를 확인했다.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 기본값 `on` | `COMMIT`이 **WAL이 디스크에 내려간 뒤** 돌아온다 |
|
||||
| **`off`** | **WAL 플러시를 기다리지 않고** 즉시 돌아온다 |
|
||||
|
||||
**결과: PostgreSQL이 갑자기 죽으면 직전 수백 밀리초의 세션 갱신이 사라질 수
|
||||
있다.** 커밋했다고 응답해놓고 없어진다.
|
||||
|
||||
Keycloak이 이걸 의도적으로 켠 이유는 명확하다 — `LAST_SESSION_REFRESH` 갱신은
|
||||
**초당 수백 번 일어나고, 몇백 밀리초쯤 잃어도 사용자가 다시 갱신하면 그만**이다.
|
||||
로그인·로그아웃 같은 것과 달리 잃어도 되는 쓰기다.
|
||||
|
||||
> **DB 복구 실험(A-2)에서 그대로 관측될 지점이다.** PostgreSQL을 정상 종료가
|
||||
> 아니라 강제 종료시키면, 마지막 몇백 밀리초의 세션 갱신이 실제로 없어져야
|
||||
> 한다. 이건 버그가 아니라 **설계된 트레이드오프**다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "show synchronous_commit" # 전역 기본값은 on
|
||||
```
|
||||
|
||||
전역 설정은 `on`이고, **Keycloak이 세션 트랜잭션에만 `SET LOCAL`로 끈다.**
|
||||
`SET LOCAL`은 그 트랜잭션이 끝나면 되돌아간다.
|
||||
|
||||
### 덤: 캐시는 읽어도 채워지지 않는다
|
||||
|
||||
```
|
||||
K1_ENTRIES_BEFORE=5.0
|
||||
REFRESH_ON_K1=200
|
||||
K1_ENTRIES_AFTER=5.0 ← 갱신을 처리하고도 그대로
|
||||
```
|
||||
|
||||
**keycloak-1은 남의 세션을 DB에서 읽어 처리하고도 캐시에 담지 않았다.**
|
||||
|
||||
0c에서 세운 모델 "각 노드는 자기가 처리한 것만 캐시한다"를 더 좁혀야 한다.
|
||||
|
||||
> 캐시에 담기는 것은 **그 노드가 로그인시켜 만든 세션**뿐이다.
|
||||
> 남의 세션은 매번 DB에서 읽는다.
|
||||
|
||||
로드밸런서가 세션을 만든 노드가 아닌 쪽으로 요청을 보내면 **매번 DB를 친다.**
|
||||
세션 어피니티(sticky session)가 정확성이 아니라 **성능** 문제인 이유가 이것이다.
|
||||
|
||||
### 덤 2: jdbc-ping 하트비트가 그대로 보인다
|
||||
|
||||
```
|
||||
01:12:37.551 pid=81369 | BEGIN
|
||||
01:12:37.551 pid=81369 | DELETE from JGROUPS_PING WHERE address=$1
|
||||
01:12:37.552 pid=81369 | INSERT INTO JGROUPS_PING (address, name, cluster_name, ip, coord, last_update, coordinated_by) values (...)
|
||||
01:12:37.553 pid=81369 | COMMIT
|
||||
```
|
||||
|
||||
**디스커버리는 별도 연결(pid=81369)에서 주기적으로 자기 행을 지우고 다시
|
||||
넣는다.** 세션 트래픽과 완전히 분리된 경로다 — 11층에서 말한 "디스커버리와
|
||||
트랜스포트는 다른 경로"가 로그에서 눈으로 확인된다.
|
||||
|
||||
---
|
||||
|
||||
## 7. 그래서 무엇이 세션을 공유하는가
|
||||
|
||||
```
|
||||
로그인 (keycloak-0)
|
||||
│
|
||||
├──▶ PostgreSQL OFFLINE_USER_SESSION ← 진실의 원천. 양쪽이 본다
|
||||
│
|
||||
└──▶ keycloak-0 로컬 캐시 ← 자기 것만. 건너가지 않는다
|
||||
|
||||
keycloak-1 이 그 세션을 물으면
|
||||
│
|
||||
└──▶ 자기 캐시에 없음 → PostgreSQL 에서 읽는다
|
||||
```
|
||||
|
||||
| 계층 | 역할 | 노드 간 공유 |
|
||||
|---|---|---|
|
||||
| **PostgreSQL** | 진실의 원천 | **여기서 일어난다** |
|
||||
| **Infinispan `sessions`** | 자기 노드가 처리한 세션의 룩어사이드 캐시 | **일어나지 않는다** |
|
||||
| **Infinispan 클러스터** | 무효화 메시지, `work` 캐시 등 | 형성은 되어 있다 |
|
||||
|
||||
이건 **Keycloak 26의 의도된 설계**다. `persistent-user-sessions`가 기본이 되면서
|
||||
DB가 진실의 원천이 됐고, 세션 캐시는 **복제할 이유가 없어졌다.** 복제를 하면
|
||||
네트워크와 메모리를 쓰면서 DB와 캐시 두 벌을 정합하게 유지해야 한다.
|
||||
|
||||
---
|
||||
|
||||
## 8. 개념
|
||||
|
||||
### 8-1. `persistent-user-sessions`
|
||||
|
||||
Keycloak 25에서 도입되고 **26에서 기본값**이 된 기능. 사용자 세션을
|
||||
Infinispan에만 두지 않고 **데이터베이스에 기록**한다.
|
||||
|
||||
| | 켜져 있을 때 (기본) | 꺼져 있을 때 (volatile) |
|
||||
|---|---|---|
|
||||
| 진실의 원천 | **PostgreSQL** | Infinispan |
|
||||
| 전체 재시작 후 | **세션이 남는다** | 전부 사라진다 |
|
||||
| 노드 간 공유 | DB가 한다 | **복제가 해야 한다** |
|
||||
| 로그인당 비용 | DB 쓰기 | 네트워크 복제 |
|
||||
|
||||
**이 실험의 결론은 전부 "켜져 있을 때"의 이야기다.** 끄면 다른 그림이 나오고,
|
||||
그 비교가 로드맵 A-2다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec keycloak-0 -- \
|
||||
/opt/keycloak/bin/kc.sh show-config 2>/dev/null | grep -i feature
|
||||
```
|
||||
|
||||
### 8-2. 온라인 세션이 `OFFLINE_` 테이블에 들어간다
|
||||
|
||||
**`USER_SESSION` 테이블은 존재하지 않는다.** 처음에 이걸 찾다가 없어서 당황했다.
|
||||
|
||||
```
|
||||
public | auth_session | table | keycloak
|
||||
public | jgroups_ping | table | keycloak
|
||||
public | offline_client_session | table | keycloak
|
||||
public | offline_user_session | table | keycloak
|
||||
public | revoked_token | table | keycloak
|
||||
public | root_auth_session | table | keycloak
|
||||
```
|
||||
|
||||
`persistent-user-sessions`는 **기존 오프라인 세션 테이블을 재사용**하고
|
||||
`offline_flag` 컬럼으로 구분한다.
|
||||
|
||||
| `offline_flag` | 의미 |
|
||||
|---|---|
|
||||
| **`'0'`** | **온라인 세션** (일반 로그인) |
|
||||
| `'1'` | 오프라인 세션 (`offline_access`) |
|
||||
|
||||
기본키가 `(user_session_id, offline_flag)` 복합키인 이유다 — 같은 세션 id가
|
||||
온라인/오프라인 두 행으로 존재할 수 있다.
|
||||
|
||||
```sql
|
||||
select offline_flag, count(*) from offline_user_session group by offline_flag;
|
||||
select user_session_id, offline_flag, created_on, last_session_refresh
|
||||
from offline_user_session where user_session_id = '<sid>';
|
||||
```
|
||||
|
||||
**이름이 내용을 배신하는 스키마다.** 운영에서 "온라인 세션이 DB 어디 있냐"를
|
||||
찾을 때 이걸 모르면 한참 헤맨다.
|
||||
|
||||
### 8-3. `sid` — 토큰과 DB를 잇는 열쇠
|
||||
|
||||
```
|
||||
JWT access_token 의 sid jiv3rVZi1VeaO07oVJkL_MYW
|
||||
↕ 같은 값
|
||||
DB user_session_id jiv3rVZi1VeaO07oVJkL_MYW
|
||||
↕ 같은 값
|
||||
Admin API 세션 목록의 id jiv3rVZi1VeaO07oVJkL_MYW
|
||||
```
|
||||
|
||||
세 곳에서 같은 문자열이다. **장애를 추적할 때 이 값 하나로 토큰·DB·관리 API를
|
||||
꿰뚫을 수 있다.** 백채널 로그아웃의 `sid` 클레임도 이것이다.
|
||||
|
||||
### 8-4. `openid` scope가 없으면 OIDC 토큰이 아니다
|
||||
|
||||
`admin-cli`에 `scope` 없이 direct grant를 하면 나오는 클레임은 이렇다.
|
||||
|
||||
```
|
||||
--- access_token ---
|
||||
클레임: azp, exp, iat, iss, jti, scope, sid, typ
|
||||
typ = Bearer | sub = None | sid = OxikTqdHCJ7ESm2GPKI1oa6c
|
||||
```
|
||||
|
||||
**`sub`이 없다.** OIDC가 아니라 순수 OAuth2 액세스 토큰이기 때문이다.
|
||||
`sub`은 OIDC가 요구하는 클레임이고, `openid` scope가 있어야 붙는다.
|
||||
|
||||
같은 이유로 `userinfo`가 403 `insufficient_scope`를 준다 — userinfo는 OIDC
|
||||
엔드포인트다. **두 현상은 하나의 원인**이다.
|
||||
|
||||
### 8-5. 룩어사이드(lookaside) 캐시
|
||||
|
||||
```
|
||||
읽기: 캐시 확인 → 없으면 DB → 캐시에 채움
|
||||
쓰기: DB 에 쓰고 → 캐시에도 씀
|
||||
```
|
||||
|
||||
캐시가 **DB 앞에 서 있되 DB를 대체하지 않는** 구조. 캐시를 통째로 날려도
|
||||
정확성은 유지되고 느려지기만 한다. Keycloak 26의 세션 캐시가 이 모양이다.
|
||||
|
||||
이 성질이 **노드 상실 실험(A-3)의 결과를 미리 결정한다** — 노드가 죽으면
|
||||
그 노드의 캐시는 사라지지만 세션은 DB에 있으므로 살아남아야 한다.
|
||||
|
||||
---
|
||||
|
||||
## 9. 다음 실험에 대한 예측
|
||||
|
||||
기준선이 생겼으므로 **틀릴 수 있는 예측**을 세울 수 있다. 예측이 빗나가면
|
||||
그것이야말로 배울 거리다.
|
||||
|
||||
| 실험 | 예측 | 근거 |
|
||||
|---|---|---|
|
||||
| **A-1** TCP 7800 차단 | **세션 공유는 안 깨진다.** 대신 무효화 전파와 `work` 캐시가 깨진다 | 세션은 7800으로 오가지 않는다 |
|
||||
| **A-2** DB 손실 | **즉시 전면 장애.** 캐시에 있는 세션도 못 쓴다 | DB가 진실의 원천 |
|
||||
| **A-2'** DB **강제** 종료 | 직전 수백 ms 의 세션 갱신이 **사라진다** | `synchronous_commit OFF` |
|
||||
| **B-5** 동시 갱신 경쟁 | 한쪽이 `VERSION` 검사에서 지고 재시도한다 | 낙관적 락 |
|
||||
| **A-3** 노드 상실 (kc-lab-2) | **세션은 살아남는다.** 죽은 노드의 캐시만 사라진다 | 룩어사이드 |
|
||||
| **A-4** volatile 비교 | 7800 차단이 **A-1과 정반대로** 치명적이 된다 | 그때는 캐시가 진실의 원천 |
|
||||
|
||||
특히 A-1은 **직관과 어긋나는 예측**이다. "클러스터 포트를 막으면 세션이
|
||||
깨진다"가 상식이지만, 이 기준선이 맞다면 안 깨져야 한다.
|
||||
|
||||
---
|
||||
|
||||
## 10. 겪은 함정
|
||||
|
||||
### 10-1. kubectl 스트림에서 출력이 통째로 사라졌다
|
||||
|
||||
`kubectl run --rm -i ... | grep` 로 받으면 **중간 조각이 유실됐다.**
|
||||
keycloak-1의 스냅샷과 그 다음 마커가 함께 없어져, 전값이 0으로 잡히면서
|
||||
**가짜 델타가 만들어졌다.**
|
||||
|
||||
```
|
||||
###BEFORE_K1 ← 여기 있어야 할 지표 20줄과
|
||||
http_code=200 다음 마커 ###LOGIN 이 통째로 사라졌다
|
||||
###AFTER_K0
|
||||
```
|
||||
|
||||
이때 리포트는 keycloak-1이 `+9`, `+7` 증가한 것처럼 보였다. **없는 복제가
|
||||
있는 것처럼 보이는, 가장 나쁜 종류의 오류다.**
|
||||
|
||||
| 고친 방법 | |
|
||||
|---|---|
|
||||
| 파드 안에서 파일로 모으고 마지막에 `cat` 한 번 | 스트리밍 중 유실을 없앤다 |
|
||||
| 스냅샷이 비면 **경고를 출력**한다 | 조용히 0으로 계산되는 것을 막는다 |
|
||||
|
||||
```bash
|
||||
for n in ('BEFORE_K0','BEFORE_K1','AFTER_K0','AFTER_K1'):
|
||||
if not blocks.get(n):
|
||||
print(f' !! {n} 스냅샷이 비었다 — 델타를 신뢰할 수 없다')
|
||||
```
|
||||
|
||||
**계측 코드는 자기가 실패했는지 스스로 말해야 한다.**
|
||||
|
||||
### 10-2. DB에서 직접 지우면 캐시는 남는다
|
||||
|
||||
정리하려고 `delete from offline_user_session`을 실행했더니, **캐시 엔트리는
|
||||
그대로 남아** 캐시 합계(19)와 DB 총계(15)가 어긋났다.
|
||||
|
||||
> 운영에서 세션 테이블을 직접 손대면 캐시와 DB가 갈라진다. 세션을 지울 때는
|
||||
> 관리 API(`logout-all`)를 쓰거나, DB를 건드렸다면 **파드를 재시작**해야 한다.
|
||||
|
||||
이 실험의 최종 수치는 **파드 재시작 후** 다시 잰 것이다.
|
||||
|
||||
### 10-3. Keycloak 이미지에는 `curl`이 없다
|
||||
|
||||
`kubectl exec keycloak-0 -- curl` 은 실패한다. 임시 `curlimages/curl` 파드를
|
||||
띄워 파드 네트워크 안에서 호출했다. 파드 IP는 클러스터 밖에서 닿지 않으므로
|
||||
이 방법이 사실상 유일하다.
|
||||
|
||||
### 10-4. 중첩 셸의 변수 치환
|
||||
|
||||
`ssh host '... $VAR ...'` 안에 다시 `sh -c "..."` 를 넣으면 인용이 세 겹이 되어
|
||||
치환이 조용히 깨진다. 첫 시도에서 파드 IP가 빈 문자열이 되어 아무 출력도
|
||||
나오지 않았다.
|
||||
|
||||
**스크립트 파일로 만들어 `scp` 로 옮기는 쪽이 옳다.** 재현도 되고 저장소에
|
||||
남는다. `deploy/lab/scripts/` 아래 세 스크립트가 그 결과다.
|
||||
|
||||
---
|
||||
|
||||
## 11. 재현
|
||||
|
||||
```bash
|
||||
# 1. 깨끗한 상태로 되돌린다 (DB 비우고 캐시 비우기)
|
||||
ssh test-server '
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "delete from offline_user_session"
|
||||
kubectl -n keycloak-lab rollout restart statefulset/keycloak
|
||||
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=300s'
|
||||
|
||||
# 2. 세 실험을 순서대로
|
||||
ssh test-server '/tmp/experiment-session-replication.sh' # 교차 노드 사용
|
||||
ssh test-server '/tmp/experiment-cache-replication-delta.sh' # 복제인가 DB인가
|
||||
ssh test-server '/tmp/experiment-cache-ownership.sh' # 엔트리 위치
|
||||
|
||||
# 3. 그래프
|
||||
# https://app2.hyeonworks.com/explore
|
||||
# vendor_statistics_approximate_entries_unique{cache="sessions"}
|
||||
# Legend: {{pod}} on {{node}}
|
||||
```
|
||||
|
||||
### 확인용 명령 모음
|
||||
|
||||
```bash
|
||||
# 클러스터 멤버
|
||||
kubectl -n keycloak-lab logs keycloak-0 | grep ISPN000094 | tail -1
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "select name, ip, coord from jgroups_ping"
|
||||
|
||||
# DB 세션
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
|
||||
-c "select offline_flag, count(*) from offline_user_session group by offline_flag"
|
||||
|
||||
# 노드별 캐시 엔트리 (파드 안에서)
|
||||
curl -s http://<pod-ip>:9000/metrics \
|
||||
| grep 'approximate_entries_unique{cache="sessions"'
|
||||
```
|
||||
@@ -0,0 +1,506 @@
|
||||
# A-1 — 노드 간 통신(TCP 7800)을 끊으면 무엇이 깨지는가
|
||||
|
||||
브랜치 `feature/keycloak-a1-jgroups-transport-block` ·
|
||||
증거 [`docs/evidence/a1-jgroups-transport-block/`](evidence/a1-jgroups-transport-block/) ·
|
||||
2026-09-04 11:38–11:52 KST · Keycloak 26.7.0 / Infinispan 16.0.12
|
||||
|
||||
맥락은 [`session-lab-prerequisites.md`](session-lab-prerequisites.md),
|
||||
기준선은 [`experiment-00-session-replication.md`](experiment-00-session-replication.md).
|
||||
|
||||
---
|
||||
|
||||
## 0. 결론부터
|
||||
|
||||
| 예측 | 결과 |
|
||||
|---|---|
|
||||
| 세션 공유는 **안 깨진다** | **맞다.** 교차 노드 refresh 가 `200` |
|
||||
| 로그아웃 전파는 **안 깨진다** | **틀렸다.** `400` 이어야 할 것이 `200` |
|
||||
| — | **NetworkPolicy 만으로는 분단이 일어나지 않는다** (예상 못 함) |
|
||||
| — | **분단된 노드가 스스로 로드밸런서에서 빠진다** (예상 못 함) |
|
||||
|
||||
**예측 하나가 빗나갔고, 예상하지 못한 것이 둘 나왔다.** 그중 하나는
|
||||
실험 방법 자체를 무효화할 뻔했다.
|
||||
|
||||
---
|
||||
|
||||
## 1. 왜 이 실험인가
|
||||
|
||||
A-0 에서 **세션은 Infinispan 복제가 아니라 PostgreSQL 로 공유된다**는 것을
|
||||
측정했다. 그렇다면 통념과 정면으로 어긋난다.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **통념** (Keycloak 24 이전 자료) | 세션은 7800 으로 복제된다 → **막으면 세션 공유가 깨진다** |
|
||||
| **A-0 측정** | 세션은 DB 로 공유된다 → **막아도 안 깨진다** |
|
||||
|
||||
둘 중 하나는 틀렸고, 이 실험이 판정한다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 기준선
|
||||
|
||||
```
|
||||
=== [기준선 1] 클러스터 뷰 ===
|
||||
keycloak-0: [keycloak-1-48749|5] (2) [keycloak-1-48749, keycloak-0-30843]
|
||||
keycloak-1: [keycloak-1-48749|5] (2) [keycloak-1-48749, keycloak-0-30843]
|
||||
|
||||
=== [기준선 2] JGROUPS_PING ===
|
||||
keycloak-0-30843 | 10.42.1.43:7800 | f
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t ← 코디네이터는 하나
|
||||
|
||||
=== [기준선 4] JGroups 지표 (양쪽 동일) ===
|
||||
fd_sock2_get_num_suspected_members 0.0
|
||||
merge3_get_num_merge_events 0.0
|
||||
nakack2_get_xmit_table_missing 0.0
|
||||
```
|
||||
|
||||
**대조군** — 차단 전에 같은 절차를 그대로 한 번 돌린다.
|
||||
|
||||
```
|
||||
=== [대조군] keycloak-0 로그인 → keycloak-1 에서 refresh ===
|
||||
sid tAWs2gCPr6SOcD4jDR9-_CzB
|
||||
keycloak-1 에서 refresh: 200
|
||||
```
|
||||
|
||||
A-0 에서 배운 규칙이다 — **시험군만 재는 측정은 측정이 아니다.**
|
||||
|
||||
---
|
||||
|
||||
## 3. 주입 — NetworkPolicy 로 7800 만 막는다
|
||||
|
||||
```bash
|
||||
kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
|
||||
```
|
||||
|
||||
```yaml
|
||||
spec:
|
||||
podSelector: { matchLabels: { app: keycloak } }
|
||||
policyTypes: [Ingress]
|
||||
ingress:
|
||||
- ports:
|
||||
- { port: 8080, protocol: TCP } # HTTP — 열어둔다
|
||||
- { port: 9000, protocol: TCP } # health+metrics — 열어둔다
|
||||
# 7800 은 일부러 없다
|
||||
```
|
||||
|
||||
### 개념 — NetworkPolicy 는 방화벽이 아니라 **허용 목록**이다
|
||||
|
||||
**"7800 을 거부"라고 쓸 수 없다.** 파드가 `policyTypes: [Ingress]` 를 가진
|
||||
정책에 선택되는 순간 **모든 인바운드가 거부**되고, 규칙에 적힌 것만 통과한다.
|
||||
그래서 7800 은 **빠뜨림으로써** 막힌다.
|
||||
|
||||
이 구조가 두 허용 규칙을 **결정적으로 만든다.** 잘못 쓰면 분단된 클러스터가
|
||||
아니라 **죽은 Keycloak 을 측정하게 된다.**
|
||||
|
||||
| 포트 | 빼면 |
|
||||
|---|---|
|
||||
| 8080 | Traefik·상대 노드의 REST 호출이 전부 끊긴다 |
|
||||
| **9000** | **readiness 프로브가 실패해 kubelet 이 파드를 죽인다** — 엉뚱한 이유로 클러스터가 깨진다 |
|
||||
|
||||
적용 직후 확인했다.
|
||||
|
||||
```
|
||||
파드 상태: keycloak-0 ready=true restarts=0
|
||||
keycloak-1 ready=true restarts=0
|
||||
9000 도달: 10.42.1.43:9000 health=200 / 10.42.0.35:9000 health=200
|
||||
8080 도달: 10.42.1.43:8080 root=200 / 10.42.0.35:8080 root=200
|
||||
```
|
||||
|
||||
**주입이 의도한 것만 건드렸음을 먼저 확인한 뒤에 결과를 해석한다.**
|
||||
|
||||
---
|
||||
|
||||
## 4. 문제 ① — **NetworkPolicy 만으로는 분단이 안 된다**
|
||||
|
||||
가장 중요한 발견이며, 하마터면 **실험 전체를 무효로 만들 뻔했다.**
|
||||
|
||||
차단 후 지표가 꿈쩍도 하지 않았다. 신규 연결은 분명히 막히는데.
|
||||
|
||||
```
|
||||
=== 7800 신규 연결 ===
|
||||
10.42.1.43:7800 curl exit=7 (연결 실패)
|
||||
10.42.1.43:9000 curl exit=28 (연결됨, telnet 이라 대기 → 타임아웃)
|
||||
```
|
||||
|
||||
그런데 파드 내부 소켓을 보니
|
||||
|
||||
```
|
||||
=== /proc/net/tcp6 · 7800 = 0x1E78 ===
|
||||
keycloak-0: ...2B012A0A:1E78 ...23002A0A:9C57 01 ← 01 = ESTABLISHED
|
||||
keycloak-1: ...23002A0A:9C57 ...2B012A0A:1E78 01
|
||||
(10.42.0.35:40023 → 10.42.1.43:7800)
|
||||
```
|
||||
|
||||
**기존 연결이 멀쩡히 살아 있다.**
|
||||
|
||||
### 왜 그런가 — conntrack
|
||||
|
||||
```
|
||||
패킷 도착
|
||||
│
|
||||
├─▶ [ conntrack: ESTABLISHED/RELATED 이면 ACCEPT ] ← 여기서 통과해버린다
|
||||
│
|
||||
└─▶ [ NetworkPolicy 규칙 평가 ] ← 여기까지 오지 않는다
|
||||
```
|
||||
|
||||
리눅스 방화벽은 성능을 위해 **이미 성립한 연결을 먼저 통과**시킨다.
|
||||
NetworkPolicy 는 그 뒤에 있으므로 **신규 연결(SYN)만** 걸러낸다.
|
||||
|
||||
```
|
||||
=== conntrack 확인 ===
|
||||
tcp 6 86398 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=40023 dport=7800 ... [ASSURED]
|
||||
tcp 6 79982 ESTABLISHED src=10.42.0.35 dst=10.42.1.43 sport=50477 dport=57800 ... [ASSURED]
|
||||
tcp 6 33 SYN_SENT src=10.42.1.58 dst=10.42.0.35 sport=34824 dport=7800 [UNREPLIED]
|
||||
─────────────────────────────────────────────────────────────
|
||||
신규 연결은 응답을 못 받는다 = 정책이 동작하고는 있다
|
||||
```
|
||||
|
||||
> **운영적 함의 — NetworkPolicy 는 이미 붙어 있는 것을 떼어내지 못한다.**
|
||||
> 보안 사고 대응으로 "지금 당장 이 통신을 끊어라"에 NetworkPolicy 를 적용하면,
|
||||
> **새 연결만 막히고 진행 중인 연결은 계속된다.** 끊으려면 conntrack 을 지우거나
|
||||
> 파드를 재시작해야 한다.
|
||||
|
||||
### 덤 — **57800 포트도 있다**
|
||||
|
||||
`sport=50477 dport=57800` — FD_SOCK2 는 **`bind_port + 50000`** 을 쓴다.
|
||||
7800 만 막고 57800 을 열어두면 장애 감지 채널이 남는다.
|
||||
이 실험의 허용 목록 방식은 **둘 다 자동으로 막았다** — 8080·9000 외 전부 거부이므로.
|
||||
|
||||
### 조치
|
||||
|
||||
```bash
|
||||
# 정확한 튜플로 지정해야 지워진다. --dport 만으로는 0건이었다
|
||||
sudo conntrack -D -p tcp -s 10.42.0.35 -d 10.42.1.43 --sport 40023 --dport 7800
|
||||
sudo conntrack -D -p tcp -s 10.42.1.43 -d 10.42.0.35 --sport 7800 --dport 40023 # 역방향
|
||||
```
|
||||
|
||||
**양쪽 노드에서, 양쪽 방향으로** 지워야 한다. 서버 쪽 노드에는 튜플이 뒤집혀
|
||||
기록되어 있다.
|
||||
|
||||
그리고 **즉시 끊기지 않는다.**
|
||||
|
||||
```
|
||||
11:41 conntrack 삭제
|
||||
11:44 cluster_size 2 → 1 ← 약 3분 뒤
|
||||
```
|
||||
|
||||
TCP 는 상대가 사라졌음을 **재전송 타임아웃**으로 알아낸다. 소켓은 한동안
|
||||
`ESTABLISHED` 로 남아 있다.
|
||||
|
||||
---
|
||||
|
||||
## 5. 문제 ② — 계측 도구가 잘못됐다
|
||||
|
||||
임시 curl 파드로 20초마다 지표를 긁었더니 이런 결과가 나왔다.
|
||||
|
||||
```
|
||||
+20초 suspected(k0 k1) = []
|
||||
+60초 suspected(k0 k1) = [0.0 0.0 0.0 0.0 ]
|
||||
+140초 suspected(k0 k1) = [0.0 ]
|
||||
```
|
||||
|
||||
**빈 값, 개수가 맞지 않는 값이 섞인다.** `kubectl run --rm` 은 매번 파드를
|
||||
만들고 지우므로 느리고 경합이 있다.
|
||||
|
||||
게다가 첫 시도의 판정 조건이
|
||||
|
||||
```sh
|
||||
[ "$R" != "0.0 0.0 " ] && echo "→ 변화 감지" && break
|
||||
```
|
||||
|
||||
여서 **빈 문자열을 "변화"로 읽고 즉시 빠져나왔다.** A-0 에서 똑같은 실수를
|
||||
했는데 또 했다.
|
||||
|
||||
> **임시 파드는 계측 도구가 아니다.** 15초마다 이미 긁고 있는 Prometheus 가
|
||||
> 그러라고 있는 것이다.
|
||||
|
||||
```bash
|
||||
kubectl -n observability port-forward svc/prometheus 19090:9090 &
|
||||
curl -s "http://localhost:19090/api/v1/query_range?query=vendor_cluster_size&start=$START&end=$END&step=60"
|
||||
```
|
||||
|
||||
그리고 이 과정에서 **`vendor_cluster_size`** 를 발견했다 — 멤버 수를 직접
|
||||
알려주는 지표다. 처음부터 이걸 봤어야 했다.
|
||||
|
||||
```bash
|
||||
curl -s "http://localhost:19090/api/v1/label/__name__/values" | grep -E "cluster|member|view"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. 진짜 분단이 일어난 순간
|
||||
|
||||
```
|
||||
=== vendor_cluster_size ===
|
||||
keycloak-1: 11:43:57=2 11:44:27=1 ... 11:51:28=2
|
||||
keycloak-0: 11:43:57=2 (파드 교체) 11:45:27=1 ... 11:51:28=2
|
||||
```
|
||||
|
||||

|
||||
|
||||
정책이 걸린 채 `keycloak-0` 이 재시작되자, 로그가 정확히 말해준다.
|
||||
|
||||
```
|
||||
GMS: JOIN(keycloak-0-26403) sent to keycloak-1-48749 timed out ← 10회
|
||||
GMS: too many JOIN attempts (10): becoming singleton ← 포기
|
||||
ISPN000094: new cluster view [keycloak-0-26403|0] (1) [keycloak-0-26403]
|
||||
```
|
||||
|
||||
`keycloak-1` 쪽도 혼자가 되었다.
|
||||
|
||||
```
|
||||
ISPN000094: [keycloak-1-48749|6] (1) [keycloak-1-48749]
|
||||
```
|
||||
|
||||
### **DB 에는 둘 다 있는데 클러스터는 안 붙는다** — 예측한 그 상태
|
||||
|
||||
```
|
||||
=== JGROUPS_PING ===
|
||||
name | ip | coord
|
||||
------------------+-----------------+-------
|
||||
keycloak-0-26403 | 10.42.1.67:7800 | t ← 코디네이터
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | t ← 코디네이터
|
||||
```
|
||||
|
||||
**`coord = t` 가 둘.** 교과서적인 split brain 이며, **데이터베이스 한 줄로
|
||||
확인된다.** 디스커버리(DB)는 살아 있고 트랜스포트(7800)만 죽은 상태다.
|
||||
|
||||
**단일 노드에서는 만들 수 없는 고장**이며, 이 실험대를 2 VM 으로 만든 이유다.
|
||||
|
||||
---
|
||||
|
||||
## 7. 본 시험 — 분단 상태에서 세션은 어떻게 되는가
|
||||
|
||||
```
|
||||
[1] keycloak-0 로그인 sid=nShl5TaBrZnKStDqaspjgmJB
|
||||
[2] keycloak-1 에서 refresh HTTP 200 ← 예측대로
|
||||
[3] keycloak-1 에서 로그아웃 HTTP 204
|
||||
[4] keycloak-0 에서 재갱신 시도 HTTP 200 ← 400 이어야 했다
|
||||
```
|
||||
|
||||
### [2] 세션 공유 — **예측이 맞았다**
|
||||
|
||||
클러스터가 갈라졌는데도 **한쪽에서 만든 세션을 반대쪽이 갱신했다.**
|
||||
A-0 의 모델이 맞고, **통념이 틀렸다.** 세션은 7800 으로 다니지 않는다.
|
||||
|
||||
### [4] 로그아웃 전파 — **예측이 틀렸다**
|
||||
|
||||
A-0 에서는 같은 절차가 `400 invalid_grant / Session not active` 였다.
|
||||
분단 상태에서는 `200` 이다. **로그아웃한 세션이 반대편에서 살아 있다.**
|
||||
|
||||
기제를 확정했다.
|
||||
|
||||
```
|
||||
=== 그 sid 가 DB 에 남아 있는가 ===
|
||||
user_session_id | offline_flag | last_session_refresh
|
||||
-----------------+--------------+----------------------
|
||||
(0 rows) ← DB 행은 삭제되었다
|
||||
|
||||
=== 노드별 세션 캐시 엔트리 ===
|
||||
keycloak-1 kc-lab-1 = 0
|
||||
keycloak-0 kc-lab-2 = 1 ← 캐시에는 남아 있다
|
||||
```
|
||||
|
||||
```
|
||||
keycloak-1 로그아웃
|
||||
│
|
||||
├──▶ PostgreSQL 행 삭제 ✔ 되었다
|
||||
│
|
||||
└──▶ keycloak-0 에게 "캐시에서 지워라" ✗ 7800 이 막혀 못 갔다
|
||||
│
|
||||
keycloak-0 은 자기 캐시로 200 을 준다 ◀────────────┘
|
||||
```
|
||||
|
||||
### **A-0 의 결론을 정정한다**
|
||||
|
||||
A-0 에서 나는 이렇게 썼다.
|
||||
|
||||
> 로그아웃과 함께 DB 행이 사라졌다 → 무효화가 DB 삭제로 전파된다
|
||||
|
||||
**그 인과는 틀렸다.** DB 행 삭제는 일어나지만, **반대편 노드는 DB 를 다시
|
||||
읽지 않는다.** 자기 캐시에 있으면 그걸로 답한다.
|
||||
|
||||
> **룩어사이드 캐시는 읽을 때 DB 와 대조하지 않는다.**
|
||||
> 캐시 무효화는 **클러스터 메시지(7800)를 타고** 간다.
|
||||
|
||||
A-0 에서 400 이 나온 것은 DB 덕분이 아니라 **그때는 7800 이 살아 있어서**였다.
|
||||
두 실험을 붙여야 비로소 정확한 그림이 나온다.
|
||||
|
||||
| | 세션 **조회** | 세션 **무효화** |
|
||||
|---|---|---|
|
||||
| 경로 | PostgreSQL | **클러스터 메시지 (7800)** |
|
||||
| 7800 차단 시 | 정상 | **전파되지 않음** |
|
||||
|
||||
---
|
||||
|
||||
## 8. 그런데 안전장치가 있었다 — 예상 못 한 발견
|
||||
|
||||
`keycloak-0` 이 `Ready=false` 였다. 이유를 물었더니
|
||||
|
||||
```json
|
||||
{ "status": "DOWN",
|
||||
"checks": [
|
||||
{ "name": "Keycloak cluster health check", "status": "DOWN",
|
||||
"data": { "Failing since": "2026-09-04 02:45:14,251" } },
|
||||
{ "name": "Keycloak database connections async health check", "status": "UP" }
|
||||
] }
|
||||
```
|
||||
|
||||
**Keycloak 은 클러스터 분단을 readiness 로 신고한다.** 그리고 쿠버네티스가
|
||||
그 신고를 받아 처리했다.
|
||||
|
||||
```
|
||||
=== Service 엔드포인트 ===
|
||||
ready 주소: [10.42.0.35] ← keycloak-1 만 트래픽을 받는다
|
||||
notReady : [10.42.1.67] ← keycloak-0 은 제외되었다
|
||||
|
||||
=== 외부 진입점 ===
|
||||
https://auth.hyeonworks.com/realms/master HTTP 200
|
||||
토큰 발급 HTTP 200
|
||||
```
|
||||
|
||||
**분단된 노드가 스스로 로드밸런서에서 빠졌고, 서비스는 계속되었다.**
|
||||
|
||||
### 그래서 7절의 로그아웃 우회는 어떻게 봐야 하나
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 내가 한 것 | Service 를 우회해 **파드 IP 로 직접** 호출 |
|
||||
| 실제 사용자 | nginx → Traefik → **Service** → Ready 인 파드만 |
|
||||
|
||||
**정문으로 들어오면 낡은 캐시에 닿지 않는다.** readiness 게이트가 막는다.
|
||||
|
||||
> 다만 이건 **비대칭이라서 살았다.** `keycloak-1` 은 원래 뷰에서 멤버가 하나
|
||||
> 줄어든 정상적인 사건이라 Ready 를 유지했고, `keycloak-0` 은 합류 자체를
|
||||
> 못 해 DOWN 이 되었다. **양쪽이 동시에 DOWN 이 되는 경로가 있다면 전면 장애다.**
|
||||
> A-5(비대칭 파티션)에서 이어서 본다.
|
||||
|
||||
---
|
||||
|
||||
## 9. 복구
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
|
||||
```
|
||||
|
||||
```
|
||||
+30초 keycloak-0=1 keycloak-1=1
|
||||
+60초 keycloak-0=1 keycloak-1=1
|
||||
+90초 keycloak-0=2 keycloak-1=2 ← 재형성
|
||||
```
|
||||
|
||||
**90초 만에 자동으로 다시 붙었다. 사람 손이 필요 없었다.**
|
||||
|
||||
```
|
||||
=== MERGE3 가 합쳤는가 ===
|
||||
merge_events keycloak-0 = 1
|
||||
merge_events keycloak-1 = 1
|
||||
```
|
||||
|
||||
**MERGE3 가 한 일이다.** split brain 을 감지해 뷰를 병합하는 프로토콜이며,
|
||||
지표가 `0 → 1` 로 올라간 것이 그 증거다.
|
||||
|
||||
```
|
||||
=== JGROUPS_PING ===
|
||||
keycloak-0-26403 | 10.42.1.67:7800 | t
|
||||
keycloak-1-48749 | 10.42.0.35:7800 | f ← 코디네이터가 하나로 돌아왔다
|
||||
```
|
||||
|
||||
**코디네이터가 keycloak-1 에서 keycloak-0 으로 넘어갔다.** 코디네이터는
|
||||
특권이 아니라 역할이며, 병합 시 재선출된다.
|
||||
|
||||
---
|
||||
|
||||
## 10. 개념 정리
|
||||
|
||||
### conntrack — 연결 추적
|
||||
|
||||
리눅스 커널이 **진행 중인 연결을 기억**하는 표. 패킷마다 규칙을 다시 평가하지
|
||||
않기 위해 존재한다.
|
||||
|
||||
| 상태 | 뜻 |
|
||||
|---|---|
|
||||
| `NEW` | 첫 패킷(SYN) |
|
||||
| **`ESTABLISHED`** | **양방향 통신이 성립함 — 규칙 평가를 건너뛴다** |
|
||||
| `[ASSURED]` | 충분히 오래된 연결. 표가 꽉 차도 안 지워진다 |
|
||||
| `SYN_SENT [UNREPLIED]` | 보냈는데 답이 없음 = **차단되고 있다** |
|
||||
|
||||
```bash
|
||||
sudo conntrack -L | grep 7800
|
||||
sudo conntrack -D -p tcp -s <src> -d <dst> --sport <sp> --dport <dp>
|
||||
```
|
||||
|
||||
### FD_SOCK2 와 포트 규약
|
||||
|
||||
| 프로토콜 | 포트 | 하는 일 |
|
||||
|---|---|---|
|
||||
| TCP (트랜스포트) | **7800** | 클러스터 메시지 |
|
||||
| **FD_SOCK2** | **57800** = 7800 + 50000 | 소켓으로 상대 생존 감시 |
|
||||
|
||||
**방화벽 규칙을 손으로 쓸 때 57800 을 빠뜨리기 쉽다.**
|
||||
|
||||
### MERGE3
|
||||
|
||||
split brain 이 생긴 뒤 **갈라진 뷰를 다시 합치는** JGroups 프로토콜.
|
||||
주기적으로 다른 코디네이터의 존재를 확인하고, 발견하면 병합을 개시한다.
|
||||
|
||||
```promql
|
||||
vendor_jgroups_merge3_get_num_merge_events
|
||||
```
|
||||
|
||||
### readiness 프로브와 Service 엔드포인트
|
||||
|
||||
```
|
||||
readiness 실패 → 파드가 Service 의 notReadyAddresses 로 이동
|
||||
→ kube-proxy 가 그 파드로 라우팅하지 않음
|
||||
→ 살아 있지만 트래픽은 안 받음
|
||||
```
|
||||
|
||||
**liveness 와 다르다.** liveness 실패는 **재시작**, readiness 실패는
|
||||
**격리**다. 클러스터 분단처럼 "재시작해도 안 나아지는" 문제에는 readiness 가
|
||||
맞는 신호다.
|
||||
|
||||
---
|
||||
|
||||
## 11. 재현 절차 (명령어)
|
||||
|
||||
```bash
|
||||
# 0. 기준선
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
|
||||
-c "select name, ip, coord from jgroups_ping order by name"
|
||||
kubectl -n observability port-forward svc/prometheus 19090:9090 &
|
||||
curl -s "http://localhost:19090/api/v1/query?query=vendor_cluster_size"
|
||||
|
||||
# 1. 차단
|
||||
kubectl apply -f deploy/lab/k8s/a1-block-jgroups-transport.yaml
|
||||
|
||||
# 2. 주입이 의도한 것만 건드렸는지 확인 (8080/9000 은 살아 있어야 한다)
|
||||
kubectl -n keycloak-lab get pods -o wide | grep keycloak # restarts=0 확인
|
||||
|
||||
# 3. 기존 연결이 남아 있음을 확인 — 이걸 안 하면 실험이 무효다
|
||||
ssh kc-lab-1 'sudo conntrack -L | grep 7800'
|
||||
|
||||
# 4. conntrack 삭제 (양쪽 노드, 양쪽 방향). 반영까지 약 3분
|
||||
ssh kc-lab-1 'sudo conntrack -D -p tcp -s <k1ip> -d <k0ip> --sport <sp> --dport 7800'
|
||||
ssh kc-lab-2 'sudo conntrack -D -p tcp -s <k0ip> -d <k1ip> --sport 7800 --dport <sp>'
|
||||
|
||||
# 5. 분단 확인
|
||||
curl -s "http://localhost:19090/api/v1/query?query=vendor_cluster_size"
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
|
||||
-c "select name, coord from jgroups_ping" # coord=t 가 둘이면 split brain
|
||||
|
||||
# 6. 복구
|
||||
kubectl -n keycloak-lab delete networkpolicy a1-block-jgroups-transport
|
||||
curl -s "http://localhost:19090/api/v1/query?query=vendor_jgroups_merge3_get_num_merge_events"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 12. 다음 실험에 남기는 것
|
||||
|
||||
| 실험 | 이 실험이 준 것 |
|
||||
|---|---|
|
||||
| **A-5** 비대칭 파티션 | **양쪽이 동시에 NotReady 가 되는 경로가 있는가.** 여기서는 비대칭이라 살았다 |
|
||||
| **A-2** DB 정지 | 캐시가 DB 와 대조하지 않는다는 사실 → **캐시에 있는 세션은 DB 없이도 읽힐 수 있다** |
|
||||
| **A-7** volatile 비교 | 같은 주입에서 세션 공유가 **깨져야** 한다. 이 실험이 그 대조군 |
|
||||
| 전체 | **주입이 실제로 걸렸는지 먼저 확인한다.** NetworkPolicy 는 기존 연결을 못 끊는다 |
|
||||
@@ -0,0 +1,308 @@
|
||||
# A-2 — PostgreSQL 이 죽으면 어떻게 되는가
|
||||
|
||||
브랜치 `feature/keycloak-a2-database-loss` ·
|
||||
증거 [`docs/evidence/a2-database-loss/`](evidence/a2-database-loss/) ·
|
||||
2026-09-04 11:56–11:58 KST · Keycloak 26.7.0
|
||||
|
||||
선행: [`A-0`](experiment-00-session-replication.md) ·
|
||||
[`A-1`](experiment-a1-jgroups-transport-block.md)
|
||||
|
||||
---
|
||||
|
||||
## 0. 결론부터
|
||||
|
||||
| 예측 | 결과 |
|
||||
|---|---|
|
||||
| 즉시 전면 장애 | **맞다.** 외부 진입점 **503**, 양쪽 노드 NotReady |
|
||||
| 캐시에 있어도 못 쓴다 | **맞다.** 캐시를 가진 노드도 `500` |
|
||||
| — | **`up = 1` 인 채로 전면 장애가 났다** (관측의 함정) |
|
||||
| — | **DB 복귀 15초 만에 재시작 없이 자동 회복** |
|
||||
|
||||
**A-1 과 정반대다.** A-1 은 한쪽만 빠지고 서비스가 계속됐지만,
|
||||
A-2 는 **살아남는 노드가 없다.**
|
||||
|
||||
---
|
||||
|
||||
## 1. 설계 — 네 경로를 구분해서 본다
|
||||
|
||||
A-1 에서 **"룩어사이드 캐시는 읽을 때 DB 와 대조하지 않는다"** 를 확인했다.
|
||||
그렇다면 캐시를 가진 노드는 DB 없이도 버틸지 모른다. 그 가설을 가른다.
|
||||
|
||||
| # | 경로 | 무엇을 보는가 |
|
||||
|---|---|---|
|
||||
| ① | **캐시를 가진 노드**에서 refresh | 캐시가 DB 를 대신할 수 있는가 |
|
||||
| ② | 캐시가 없는 노드에서 refresh | 완전한 DB 의존 |
|
||||
| ③ | 새 로그인 | 쓰기 경로 |
|
||||
| ④ | 이미 발급된 토큰으로 조회 | 서명만으로 되는 경로 |
|
||||
|
||||
**access token 수명이 60초**이므로, 토큰 발급 → DB 정지 → 시험을 그 안에
|
||||
끝내야 한다.
|
||||
|
||||
### 계측 도구를 바꿨다
|
||||
|
||||
A-1 에서 임시 curl 파드가 형편없는 계측 도구임을 확인했다. 여기서는
|
||||
**상주 탐침 파드**를 하나 띄우고 `exec` 로 단계를 이어간다. 토큰을 파드 안
|
||||
파일에 남겨 **DB 정지 전후로 같은 토큰**을 쓸 수 있다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab run a2-probe --image=curlimages/curl:8.11.1 \
|
||||
--restart=Never --command -- sleep 7200
|
||||
kubectl -n keycloak-lab wait --for=condition=Ready pod/a2-probe --timeout=120s
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 기준선
|
||||
|
||||
```
|
||||
keycloak-0 ready=true 10.42.1.67 kc-lab-2
|
||||
keycloak-1 ready=true 10.42.0.35 kc-lab-1
|
||||
postgres ready=true 10.42.1.24 kc-lab-2
|
||||
|
||||
cluster_size keycloak-0 = 2
|
||||
cluster_size keycloak-1 = 2
|
||||
```
|
||||
|
||||
세션을 양쪽에 하나씩 만들고, A-0 대로 **각자 자기 노드에만 캐시**되는 것을
|
||||
확인했다.
|
||||
|
||||
```
|
||||
keycloak-0 에서 로그인 sid=EAXV5HcG2J1BZ3vnwONf64AQ
|
||||
keycloak-1 에서 로그인 sid=McyTj5lj3n_JqApCXeuAHExc
|
||||
→ 캐시 keycloak-0 = 1 건 / keycloak-1 = 0 건 (스크레이프 지연)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 주입
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
|
||||
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=90s
|
||||
```
|
||||
|
||||
```
|
||||
정지 시각: 11:56:04
|
||||
삭제 완료: 11:56:04 ← 즉시
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 결과 — 네 경로
|
||||
|
||||
```
|
||||
① 캐시를 가진 노드(keycloak-0)에서 refresh HTTP 500
|
||||
② 캐시가 없는 노드(keycloak-1)에서 refresh HTTP 500
|
||||
③ 새 로그인 HTTP 500
|
||||
④ 관리 API (세션 조회 필요) HTTP 500
|
||||
|
||||
--- 오류 본문 ---
|
||||
{"error":"unknown_error","error_description":"For more on this error consult the server log."}
|
||||
```
|
||||
|
||||
### ① 이 500 인 것이 중요하다
|
||||
|
||||
**캐시에 세션을 들고 있어도 refresh 는 실패한다.**
|
||||
|
||||
A-1 에서는 로그아웃된 세션을 캐시로 `200` 을 줬다. 왜 여기서는 안 되는가.
|
||||
|
||||
```
|
||||
refresh 처리
|
||||
├── 세션이 존재하는가 → 캐시로 답할 수 있다
|
||||
└── LAST_SESSION_REFRESH 갱신 → DB 쓰기가 필요하다 ← 여기서 죽는다
|
||||
```
|
||||
|
||||
A-0 에서 잡은 SQL 그대로다.
|
||||
|
||||
```sql
|
||||
update OFFLINE_USER_SESSION set LAST_SESSION_REFRESH=$1, VERSION=$2 where ...
|
||||
```
|
||||
|
||||
> **캐시는 읽기를 대신할 뿐, 쓰기를 대신하지 못한다.**
|
||||
> refresh 는 이름과 달리 **쓰기 연산**이다.
|
||||
|
||||
### 로그가 말하는 원인
|
||||
|
||||
```
|
||||
Caused by: java.net.ConnectException: Connection refused
|
||||
at org.postgresql.core.v3.ConnectionFactoryImpl.tryConnect
|
||||
at io.agroal.pool.ConnectionPool$CreateConnectionTask.call
|
||||
```
|
||||
|
||||
`agroal` 은 Quarkus 의 커넥션 풀이다. 풀이 새 커넥션을 만들지 못한다.
|
||||
|
||||
---
|
||||
|
||||
## 5. 살아남은 것 — 상태가 필요 없는 경로
|
||||
|
||||
```
|
||||
JWKS 엔드포인트(realm 공개키) HTTP 200
|
||||
realm 메타데이터(.well-known) HTTP 200
|
||||
관리 API (세션 조회 필요) HTTP 500
|
||||
```
|
||||
|
||||
**realm 공개키와 메타데이터는 메모리에 있으므로 DB 없이도 응답한다.**
|
||||
|
||||
이론적으로는 **이미 JWKS 를 캐시한 리소스 서버는 토큰 검증을 계속할 수 있다**는
|
||||
뜻이다. 다만 이 실험대에는 독립 리소스 서버가 아직 없으므로 **여기까지가
|
||||
말할 수 있는 범위**다 — B층에서 확인한다.
|
||||
|
||||
> **그런데 정문으로는 이것도 못 쓴다.** 아래 6절 때문이다.
|
||||
|
||||
---
|
||||
|
||||
## 6. 전면 장애 — 살아남는 노드가 없다
|
||||
|
||||
```
|
||||
=== 파드 Ready ===
|
||||
keycloak-0 false restarts=0
|
||||
keycloak-1 false restarts=0
|
||||
|
||||
=== Service 엔드포인트 ===
|
||||
ready : [] ← 비었다
|
||||
notReady: [10.42.0.35 10.42.1.67]
|
||||
|
||||
=== 외부 진입점 ===
|
||||
https://auth.hyeonworks.com/realms/master HTTP 503
|
||||
```
|
||||
|
||||
```json
|
||||
{ "status": "DOWN",
|
||||
"checks": [
|
||||
{ "name": "Keycloak cluster health check", "status": "UP" },
|
||||
{ "name": "Keycloak database connections async health check", "status": "DOWN" },
|
||||
{ "name": "Keycloak Initialized", "status": "UP" } ] }
|
||||
```
|
||||
|
||||
**`cluster health` 는 UP 인데 `database connections` 가 DOWN 이라 전체가 DOWN 이다.**
|
||||
헬스체크는 **모든 항목이 UP 이어야 UP** 이다.
|
||||
|
||||
### A-1 과의 대비가 이 실험의 핵심이다
|
||||
|
||||
| | A-1 (7800 차단) | **A-2 (DB 정지)** |
|
||||
|---|---|---|
|
||||
| Ready 인 파드 | keycloak-1 **1개 생존** | **0개** |
|
||||
| Service `ready` | `[10.42.0.35]` | **`[]`** |
|
||||
| 외부 응답 | **200** | **503** |
|
||||
| 성격 | 용량 저하 | **전면 장애** |
|
||||
|
||||
**노드를 몇 대로 늘려도 DB 가 죽으면 전부 같이 죽는다.**
|
||||
Keycloak 의 대수는 DB 장애에 아무 도움이 되지 않는다.
|
||||
|
||||
> 원래 질문 *"Redis 또는 DB가 뒤질 경우 어떻게 복구를 해야 되는지"* 에 대한
|
||||
> 첫 번째 답 — **복구 이전에, DB 이중화가 Keycloak 대수보다 우선한다.**
|
||||
|
||||
---
|
||||
|
||||
## 7. 관측의 함정 — `up = 1` 인 채로 전면 장애
|
||||
|
||||
```
|
||||
up{pod=keycloak-1} = 1
|
||||
up{pod=keycloak-0} = 1 ← 서비스는 503 인데
|
||||
```
|
||||
|
||||

|
||||
|
||||
**전 구간 평평하다.** (11:44 의 짧은 골은 A-1 에서 파드를 교체한 자국이다.)
|
||||
|
||||
`up` 은 **Prometheus 가 `/metrics` 를 긁는 데 성공했는가**만 말한다.
|
||||
프로세스는 멀쩡히 살아 메트릭을 내놓고 있었다. **기능은 전멸했는데.**
|
||||
|
||||
| 지표 | 이 장애에서 |
|
||||
|---|---|
|
||||
| `up` | **1 — 아무것도 알려주지 않는다** |
|
||||
| 파드 `Ready` | **false — 여기서 드러난다** |
|
||||
| 외부 HTTP 코드 | **503 — 사용자가 겪는 것** |
|
||||
|
||||
> **A-0 에서 나는 `up` 을 "가장 중요한 합성 지표"라고 썼다.**
|
||||
> 절반만 맞다. `up` 은 **대상이 사라진 것**을 잡지만 **대상이 살아서 못 쓰는 것**은
|
||||
> 못 잡는다. 후자가 운영에서 훨씬 흔하다.
|
||||
>
|
||||
> **알림은 `up` 이 아니라 readiness 와 외부 응답 코드에 걸어야 한다.**
|
||||
|
||||
이 실험대에는 아직 `kube-state-metrics` 가 없어 파드 readiness 가 지표로
|
||||
남지 않는다. **관측 스택에 빠진 것을 이 실험이 찾아냈다** — 보완 항목이다.
|
||||
|
||||
---
|
||||
|
||||
## 8. 복구 — 자동이었다
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
|
||||
```
|
||||
|
||||
```
|
||||
재기동 시각: 11:57:09
|
||||
+15초 keycloak-0 true keycloak-1 true | 외부 HTTP 200
|
||||
→ 서비스 복귀
|
||||
|
||||
재시작 횟수: keycloak-0 = 0, keycloak-1 = 0
|
||||
정지 전 세션: online 세션 5 건 살아남음
|
||||
```
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 회복 시간 | **약 15초** (DB Ready 이후) |
|
||||
| 사람 개입 | **없음** |
|
||||
| Keycloak 재시작 | **불필요** — `restarts=0` |
|
||||
| 세션 | **살아남음** — DB 에 있으므로 |
|
||||
|
||||
**커넥션 풀이 스스로 재연결하고 readiness 가 다시 UP 이 되면서 Service 에
|
||||
복귀했다.** `readiness` 를 쓴 설계의 이득이 여기서 나온다 — `liveness` 였다면
|
||||
파드가 재시작되어 캐시까지 날아갔을 것이다.
|
||||
|
||||
### 개념 — readiness 와 liveness 를 가르는 기준
|
||||
|
||||
| | 실패하면 | 언제 쓰나 |
|
||||
|---|---|---|
|
||||
| **liveness** | **재시작** | 재시작하면 나아지는 문제 (교착, 메모리 누수) |
|
||||
| **readiness** | **트래픽에서 격리** | 재시작해도 안 나아지는 문제 (**의존 대상이 죽음**) |
|
||||
|
||||
**DB 장애에 liveness 를 걸면 재앙이다.** 모든 파드가 무한 재시작하고,
|
||||
DB 가 돌아와도 CrashLoopBackOff 의 백오프 때문에 회복이 늦어진다.
|
||||
|
||||
---
|
||||
|
||||
## 9. 재현 절차 (명령어)
|
||||
|
||||
```bash
|
||||
# 0. 상주 탐침 (임시 파드는 계측에 부적합 — A-1 참조)
|
||||
kubectl -n keycloak-lab run a2-probe --image=curlimages/curl:8.11.1 \
|
||||
--restart=Never --command -- sleep 7200
|
||||
kubectl -n keycloak-lab wait --for=condition=Ready pod/a2-probe --timeout=120s
|
||||
|
||||
# 1. 토큰 발급 (access 60초 안에 시험을 끝내야 한다)
|
||||
kubectl -n keycloak-lab exec a2-probe -- sh -c \
|
||||
'curl -s -X POST http://<k0>:8080/realms/master/protocol/openid-connect/token \
|
||||
-d grant_type=password -d client_id=admin-cli \
|
||||
-d username=admin -d password=<pw> > /tmp/tok.json'
|
||||
|
||||
# 2. DB 정지
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
|
||||
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=90s
|
||||
|
||||
# 3. 네 경로
|
||||
kubectl -n keycloak-lab exec a2-probe -- curl -s -o /dev/null -w '%{http_code}\n' ...
|
||||
|
||||
# 4. 영향 범위
|
||||
kubectl -n keycloak-lab get endpoints keycloak \
|
||||
-o jsonpath='{.subsets[*].addresses[*].ip}' # 비어 있으면 전면 장애
|
||||
curl -s -o /dev/null -w '%{http_code}\n' https://auth.hyeonworks.com/realms/master
|
||||
|
||||
# 5. up 이 거짓말하는 것을 확인
|
||||
curl -s "http://localhost:19090/api/v1/query?query=up%7Bjob=%22keycloak%22%7D"
|
||||
|
||||
# 6. 복구
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. 다음 실험에 남기는 것
|
||||
|
||||
| 실험 | 이 실험이 준 것 |
|
||||
|---|---|
|
||||
| **A-3** DB 강제 종료 | 정상 정지는 데이터를 안 잃었다. **강제 종료는?** (`synchronous_commit OFF`) |
|
||||
| **A-4** 노드 상실 | postgres 가 kc-lab-2 에 있으므로 그 노드를 죽이면 **A-2 가 함께 일어난다** |
|
||||
| **D-1** 백업·복구 | 여기서는 DB 가 되살아났다. **데이터가 사라졌다면?** |
|
||||
| 관측 스택 | **`kube-state-metrics` 가 없어 파드 readiness 가 지표로 안 남는다** — 보완 필요 |
|
||||
@@ -0,0 +1,323 @@
|
||||
# A-3 — DB 를 강제로 죽이면 무엇을 잃는가 (RPO)
|
||||
|
||||
브랜치 `feature/keycloak-a3-database-crash` ·
|
||||
증거 [`docs/evidence/a3-database-crash/`](evidence/a3-database-crash/) ·
|
||||
2026-09-04 12:00–12:05 KST · Keycloak 26.7.0 / PostgreSQL 16
|
||||
|
||||
선행: [`A-0`](experiment-00-session-replication.md) ·
|
||||
[`A-2`](experiment-a2-database-loss.md)
|
||||
|
||||
---
|
||||
|
||||
## 0. 결론부터
|
||||
|
||||
```
|
||||
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
|
||||
그중 DB 에 실제로 존재 : 149 건
|
||||
★ 유실 : 4 건
|
||||
```
|
||||
|
||||
**로그인이 성공했다고 응답받았는데 세션이 존재하지 않는다.**
|
||||
|
||||
A-0 에서 발견한 `SET LOCAL synchronous_commit TO OFF` 의 대가를 실측했다.
|
||||
버그가 아니라 **의도된 설계**이며, 그 비용이 얼마인지를 숫자로 확인한 것이다.
|
||||
|
||||
---
|
||||
|
||||
## 1. 설계 — 무엇을 재야 손실이 보이는가
|
||||
|
||||
### 1-1. `LAST_SESSION_REFRESH` 로는 못 잰다
|
||||
|
||||
처음 계획은 "세션 갱신 시각이 되감기는지" 보는 것이었다. 스키마를 보고 접었다.
|
||||
|
||||
```
|
||||
created_on | integer
|
||||
last_session_refresh | integer ← 초 단위
|
||||
```
|
||||
|
||||
**손실 창은 수백 밀리초**인데 눈금이 **1초**다. 보일 리가 없다.
|
||||
|
||||
### 1-2. 행 존재 여부로 잰다 — 이진 판정
|
||||
|
||||
```
|
||||
로그인 1회 = OFFLINE_USER_SESSION 행 1개
|
||||
클라이언트가 sid 를 받았다 = 서버가 COMMIT 했다고 응답했다
|
||||
크래시 후 그 sid 가 없다 = 잃은 것
|
||||
```
|
||||
|
||||
**있거나 없거나**이므로 눈금 문제가 없다.
|
||||
|
||||
### 1-3. 그런데 로그인도 비동기 커밋인가 — **먼저 확인해야 한다**
|
||||
|
||||
A-0 에서 잡은 것은 **refresh** 트랜잭션이었다. 로그인(INSERT)도 그런지는
|
||||
확인하지 않았다. 아니라면 이 측정 설계 자체가 성립하지 않는다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "alter system set log_statement='all'"
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
psql -U keycloak -d keycloak -c "select pg_reload_conf()"
|
||||
```
|
||||
|
||||
```
|
||||
BEGIN
|
||||
insert into OFFLINE_USER_SESSION (...) values (...)
|
||||
insert into OFFLINE_CLIENT_SESSION (...) values (...)
|
||||
SET LOCAL synchronous_commit TO OFF ← 로그인도 비동기 커밋이다
|
||||
COMMIT
|
||||
```
|
||||
|
||||
**확인됐고, 함의가 refresh 보다 훨씬 무겁다.**
|
||||
|
||||
| | 잃으면 |
|
||||
|---|---|
|
||||
| refresh 갱신 시각 | 세션 수명이 조금 짧아진다. 사용자는 모른다 |
|
||||
| **로그인 자체** | **토큰은 손에 있는데 세션이 없다.** 다음 요청부터 실패 |
|
||||
|
||||
---
|
||||
|
||||
## 2. 실패한 주입 ① — `--grace-period=0 --force` 는 크래시가 아니다
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab delete pod -l app=postgres --grace-period=0 --force
|
||||
```
|
||||
|
||||
```
|
||||
클라이언트 성공: 291 건
|
||||
DB 에 존재: 291 건
|
||||
★ 유실: 0 건
|
||||
```
|
||||
|
||||
**0건.** 그런데 이건 "안 잃었다"가 아니라 **죽인 적이 없는 것**이다.
|
||||
|
||||
```
|
||||
=== 재기동 로그 ===
|
||||
database system is ready to accept connections
|
||||
(그뿐. "not properly shut down" 이 없다)
|
||||
```
|
||||
|
||||
**crash recovery 가 돌지 않았다 = 깨끗하게 내려갔다.**
|
||||
|
||||
| 신호 | PostgreSQL 의 반응 |
|
||||
|---|---|
|
||||
| **SIGTERM** | **fast shutdown** — 진행 중 트랜잭션을 롤백하고 **WAL 을 플러시**한 뒤 종료 |
|
||||
| SIGINT | smart shutdown — 연결이 끊기길 기다린다 |
|
||||
| **SIGKILL** | **즉사** — 플러시 없음. 다음 기동에 crash recovery |
|
||||
|
||||
`--force --grace-period=0` 는 API 오브젝트를 즉시 지우지만 컨테이너 런타임은
|
||||
여전히 정상 종료 절차를 밟는다. **PostgreSQL 은 SIGTERM 을 받고 얌전히
|
||||
플러시했다.**
|
||||
|
||||
> **A-1 에서 배운 것이 또 나왔다** — 주입이 실제로 걸렸는지 먼저 확인하지
|
||||
> 않으면 **"아무 일도 없었다"를 결과로 착각한다.**
|
||||
> 여기서는 **crash recovery 메시지가 그 확인 수단**이다.
|
||||
|
||||
---
|
||||
|
||||
## 3. 실패한 주입 ② — 컨테이너 안에서 PID 1 은 SIGKILL 을 받지 않는다
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- kill -9 1
|
||||
```
|
||||
|
||||
**아무 일도 일어나지 않았다.** 파드는 재시작하지 않았고 로그 시각도 그대로였다.
|
||||
|
||||
### 개념 — PID 1 의 시그널 보호
|
||||
|
||||
리눅스 커널은 **PID 1 을 특별 취급**한다. 자기 PID 네임스페이스 안에서 온
|
||||
시그널은 **핸들러가 등록된 것만** 전달된다. **SIGKILL 도 예외가 아니다.**
|
||||
|
||||
```
|
||||
같은 네임스페이스 안에서 → PID 1 은 등록하지 않은 시그널을 무시한다
|
||||
조상 네임스페이스에서 → 전달된다 (노드에서 kill -9 하면 죽는다)
|
||||
```
|
||||
|
||||
부팅 초기에 init 을 실수로 죽여 시스템이 멈추는 것을 막기 위한 장치인데,
|
||||
컨테이너에서는 **"안에서는 PID 1 을 못 죽인다"** 로 나타난다.
|
||||
|
||||
---
|
||||
|
||||
## 4. 성공한 주입 — 백엔드 프로세스를 죽인다
|
||||
|
||||
PostgreSQL 은 **postmaster(부모) + 연결마다 백엔드(자식)** 구조다.
|
||||
자식 하나가 비정상 종료하면 **postmaster 는 공유 메모리가 오염됐다고 보고
|
||||
전체를 재초기화**한다. 그게 곧 crash recovery 다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
sh -c 'kill -9 $(pgrep -f "postgres: keycloak keycloak" | head -1)'
|
||||
```
|
||||
|
||||
```
|
||||
server process (PID 40) was terminated by signal 9: Killed
|
||||
terminating any other active server processes
|
||||
all server processes terminated; reinitializing
|
||||
database system was not properly shut down; automatic recovery in progress
|
||||
redo starts at 0/23CAB68
|
||||
redo done at 0/2529E40
|
||||
checkpoint complete: wrote 113 buffers ...
|
||||
database system is ready to accept connections
|
||||
```
|
||||
|
||||
**이번엔 주입이 걸렸다.** `not properly shut down` + `redo` 가 증거다.
|
||||
|
||||
파드는 재시작하지 않는다 (`restarts=0`) — 컨테이너의 PID 1 인 postmaster 는
|
||||
살아 있고, 자식만 갈아치운 것이다. **데이터 관점에서는 전원이 나간 것과 같다.**
|
||||
|
||||
---
|
||||
|
||||
## 5. 결과
|
||||
|
||||
```
|
||||
=== 크래시 전후 대조 ===
|
||||
클라이언트가 200 과 토큰을 받은 로그인 : 153 건
|
||||
그중 DB 에 실제로 존재 : 149 건
|
||||
★ 유실 : 4 건
|
||||
|
||||
=== 유실된 sid ===
|
||||
★ CQUfg9HLH29xvhiu6pVlfWOo ← 토큰은 발급됐는데 세션이 없다
|
||||
★ 5gLP4fqmpZBbjhH_d-0TPMMr
|
||||
★ hkcOv1QskUFmYveMLB6Hljra
|
||||
★ p5XybeQIYmAs818gO4Vl_5ea
|
||||
```
|
||||
|
||||
**약 2.6% 유실.** 초당 19건 정도 로그인하던 중이었으므로
|
||||
**대략 마지막 0.2초 분량**이다 — `wal_writer_delay` 기본값(200ms)과 맞는다.
|
||||
|
||||
### 사용자에게 어떻게 보이는가
|
||||
|
||||
```
|
||||
로그인 성공 → access token + refresh token 을 받음
|
||||
│
|
||||
│ (크래시)
|
||||
▼
|
||||
다음 요청 → access token 은 60초간 통한다
|
||||
│ (서명만 보는 경로라면)
|
||||
▼
|
||||
60초 후 refresh → "Session not active" → 다시 로그인
|
||||
```
|
||||
|
||||
**즉시 드러나지 않는다.** access token 수명 동안은 정상으로 보이다가
|
||||
갱신 시점에 끊긴다. 장애와 증상 사이에 **최대 60초의 시차**가 있다.
|
||||
|
||||
---
|
||||
|
||||
## 6. 개념
|
||||
|
||||
### WAL 과 `synchronous_commit`
|
||||
|
||||
```
|
||||
COMMIT
|
||||
│
|
||||
├─ WAL 버퍼(메모리)에 기록 ← 항상 한다
|
||||
│
|
||||
├─ synchronous_commit = on : 디스크 플러시를 기다렸다가 응답
|
||||
└─ synchronous_commit = off : 기다리지 않고 즉시 응답 ← Keycloak
|
||||
│
|
||||
└─ 크래시 시 이 구간이 사라진다
|
||||
```
|
||||
|
||||
| 설정 | 응답 속도 | 잃는 것 |
|
||||
|---|---|---|
|
||||
| `on` (PostgreSQL 기본) | 느리다 (디스크 대기) | 없다 |
|
||||
| **`off`** | 빠르다 | **최대 `wal_writer_delay` × 3 분량** |
|
||||
|
||||
**전역 설정은 `on` 이었다.**
|
||||
|
||||
```
|
||||
전역 synchronous_commit: on
|
||||
```
|
||||
|
||||
**Keycloak 이 자기 트랜잭션에만 `SET LOCAL` 로 끈다.** DBA 가 서버 설정만
|
||||
보고 "우리는 동기 커밋"이라 믿으면 틀린다. **애플리케이션이 트랜잭션 단위로
|
||||
뒤집을 수 있다.**
|
||||
|
||||
### crash recovery
|
||||
|
||||
```
|
||||
기동 시 pg_control 을 읽는다
|
||||
└─ "깨끗하게 종료됨" 표시가 없다
|
||||
└─ "database system was not properly shut down"
|
||||
└─ 마지막 체크포인트부터 WAL 을 재생(redo)
|
||||
└─ 디스크에 안 내려간 커밋은 복구할 수 없다 ← 손실
|
||||
```
|
||||
|
||||
`redo starts at 0/23CAB68` → `redo done at 0/2529E40` 사이가 재생된 구간이다.
|
||||
**WAL 에 없는 것은 재생할 수도 없다.**
|
||||
|
||||
### 이 손실이 "허용된" 이유
|
||||
|
||||
Keycloak 의 판단은 이렇게 읽힌다.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 세션 쓰기는 **매우 잦다** | 로그인마다, refresh 마다 |
|
||||
| 잃어도 **회복 가능하다** | 사용자가 다시 로그인하면 된다 |
|
||||
| 동기 커밋의 비용은 **모든 요청에 붙는다** | 크래시는 드물다 |
|
||||
|
||||
**드문 사고의 비용을 상시 지연으로 지불하지 않겠다는 선택**이다.
|
||||
합리적이지만, **선택했다는 사실을 알고 있어야 한다.**
|
||||
|
||||
---
|
||||
|
||||
## 7. 운영에 주는 것
|
||||
|
||||
| 알게 된 것 | 함의 |
|
||||
|---|---|
|
||||
| 로그인도 비동기 커밋 | **RPO 가 0 이 아니다.** 크래시 시 마지막 수백 ms 로그인은 사라진다 |
|
||||
| 전역 `on` 인데 세션만 `off` | **서버 설정으로 판단하면 안 된다.** 애플리케이션이 뒤집는다 |
|
||||
| 손실이 즉시 안 보인다 | access token 수명만큼 시차. **모니터링은 갱신 실패율을 봐야 한다** |
|
||||
| `--grace-period=0` 은 크래시가 아니다 | **장애 훈련이 훈련이 안 될 수 있다** |
|
||||
| 컨테이너 안에서 PID 1 을 못 죽인다 | 크래시 재현은 **자식 프로세스**나 **노드에서** |
|
||||
|
||||
### 바꿀 수 있는가
|
||||
|
||||
```sql
|
||||
-- 세션 트랜잭션까지 동기 커밋으로 강제하려면 (지연 대가를 치른다)
|
||||
ALTER DATABASE keycloak SET synchronous_commit = on; -- SET LOCAL 이 이깁니다
|
||||
```
|
||||
|
||||
**`SET LOCAL` 이 우선하므로 이것으로는 못 막는다.** Keycloak 설정이나
|
||||
소스 수준의 문제이며, **RPO 0 이 필요하면 복제(streaming replication)로
|
||||
푸는 것이 맞다** — 동기 스탠바이가 있으면 `synchronous_commit` 의 의미가
|
||||
달라진다.
|
||||
|
||||
---
|
||||
|
||||
## 8. 재현 절차 (명령어)
|
||||
|
||||
```bash
|
||||
# 0. 설계 확인 — 로그인도 비동기 커밋인지 먼저 본다
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
|
||||
-c "alter system set log_statement='all'" -c "select pg_reload_conf()"
|
||||
# → 로그인 1회 후 로그에서 "SET LOCAL synchronous_commit TO OFF" 확인
|
||||
|
||||
# 1. 세션 테이블 비우기
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak \
|
||||
-c "delete from offline_user_session"
|
||||
|
||||
# 2. 로그인 루프 (호스트에서 백그라운드 exec — 파드 안 & 는 exec 종료와 함께 죽는다)
|
||||
kubectl -n keycloak-lab exec a2-probe -- sh -c '<로그인 반복, sid 를 /tmp/sids 에>' &
|
||||
|
||||
# 3. 진짜 크래시 — 백엔드 프로세스에 SIGKILL
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- \
|
||||
sh -c 'kill -9 $(pgrep -f "postgres: keycloak keycloak" | head -1)'
|
||||
|
||||
# 4. 주입이 걸렸는지 확인 — 이게 없으면 결과를 해석하지 않는다
|
||||
kubectl -n keycloak-lab logs deploy/postgres | grep -E "not properly shut down|redo"
|
||||
|
||||
# 5. 대조
|
||||
kubectl -n keycloak-lab exec deploy/postgres -- psql -U keycloak -d keycloak -tAc \
|
||||
"select count(*) from offline_user_session where user_session_id in (<sid 목록>)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. 다음 실험에 남기는 것
|
||||
|
||||
| 실험 | 이 실험이 준 것 |
|
||||
|---|---|
|
||||
| **D-1** 백업·복구 | RPO 는 **백업 주기 + 이 손실**이다. 둘을 더해야 진짜 RPO |
|
||||
| **B-6** Redis 영속화 | `appendfsync everysec` 은 **같은 모양의 트레이드오프** |
|
||||
| **A-4** 노드 상실 | 노드가 죽으면 이것도 함께 일어난다 (postgres 가 kc-lab-2) |
|
||||
| 전체 | **주입 성공 신호를 미리 정한다.** 여기서는 crash recovery 로그 |
|
||||
@@ -0,0 +1,397 @@
|
||||
# A-4 — 노드가 통째로 죽으면 어떻게 되는가
|
||||
|
||||
브랜치 `feature/keycloak-a4-node-loss` ·
|
||||
증거 [`docs/evidence/a4-node-loss/`](evidence/a4-node-loss/) ·
|
||||
2026-09-04 12:07–12:24 KST
|
||||
|
||||
선행: [`A-2`](experiment-a2-database-loss.md) · [`A-3`](experiment-a3-database-crash.md)
|
||||
|
||||
**두 판본으로 나눠 실행했다.**
|
||||
|
||||
| | 죽인 노드 | 성격 |
|
||||
|---|---|---|
|
||||
| **4a** | `kc-lab-2` (워커) | Keycloak + **DB 동시 상실** |
|
||||
| **4b** | `kc-lab-1` (k3s 서버) | **컨트롤 플레인 + 진입점 상실** |
|
||||
|
||||
`virsh destroy` 는 종료 신호를 보내지 않는다. **전원을 뽑는 것과 같다.**
|
||||
|
||||
---
|
||||
|
||||
## 0. 결론부터
|
||||
|
||||
| | 4a 워커 상실 | 4b 컨트롤 플레인 상실 |
|
||||
|---|---|---|
|
||||
| 외부 응답 | **503** | **000** (연결 자체가 안 됨) |
|
||||
| `kubectl` | 정상 | **불통** |
|
||||
| 살아 있는 워크로드 | keycloak-1 (하지만 DB 없음) | **keycloak-0 은 계속 돌고 있다** |
|
||||
| 복구 시간 | **60초** | **60초** |
|
||||
|
||||
**둘 다 전면 장애**지만 이유가 다르다. 4a 는 **DB 가 같이 죽어서**, 4b 는
|
||||
**들어갈 길이 없어서**다.
|
||||
|
||||
그리고 예상하지 못한 것 셋 —
|
||||
**죽은 파드가 산 파드보다 건강해 보이고**, **StatefulSet 은 대체 파드를 만들지
|
||||
않으며**, **PVC 때문에 재배치가 불가능하다.**
|
||||
|
||||
---
|
||||
|
||||
## 4a. 워커 노드 상실 (`kc-lab-2`)
|
||||
|
||||
### 기준선
|
||||
|
||||
```
|
||||
kc-lab-1 Ready control-plane=true
|
||||
kc-lab-2 Ready <none>
|
||||
|
||||
keycloak-0 ready=true kc-lab-2
|
||||
keycloak-1 ready=true kc-lab-1
|
||||
postgres ready=true kc-lab-2
|
||||
|
||||
PVC postgres-data → kc-lab-2 ← 재배치 가능성을 여기서 이미 알 수 있다
|
||||
외부 HTTP 200
|
||||
```
|
||||
|
||||
### 주입
|
||||
|
||||
```bash
|
||||
virsh destroy kc-lab-2
|
||||
```
|
||||
|
||||
### 관찰
|
||||
|
||||
```
|
||||
+15초 node=Ready | keycloak-0=Running postgres=Running | 외부 HTTP 000
|
||||
+30초 node=Ready | keycloak-0=Running postgres=Running | 외부 HTTP 000
|
||||
+45초 node=NotReady | keycloak-0=Running postgres=Running | 외부 HTTP 503
|
||||
...
|
||||
+180초 node=NotReady | keycloak-0=Running postgres=Running | 외부 HTTP 503
|
||||
```
|
||||
|
||||
### 발견 ① — 쿠버네티스가 알아채는 데 **40초** 걸린다
|
||||
|
||||
`+15초`, `+30초` 에도 노드는 여전히 `Ready` 다. **기계는 이미 없는데.**
|
||||
|
||||
kube-controller-manager 의 `node-monitor-grace-period` (기본 40초) 동안
|
||||
kubelet 의 하트비트가 없어야 `NotReady` 로 바꾼다.
|
||||
|
||||
> **그 40초 동안 쿠버네티스는 거짓말을 한다.** 사용자는 이미 장애를 겪고 있다
|
||||
> (`000`). **노드 상태를 알림 근거로 삼으면 항상 늦는다.**
|
||||
|
||||
### 발견 ② — **죽은 파드가 산 파드보다 건강해 보인다**
|
||||
|
||||
```
|
||||
POD PHASE READY NODE
|
||||
keycloak-0 Running true kc-lab-2 ← 기계가 꺼져 있다
|
||||
keycloak-1 Running false kc-lab-1 ← 살아 있다
|
||||
```
|
||||
|
||||
**`keycloak-0` 은 `ready=true`, `keycloak-1` 은 `ready=false`.**
|
||||
|
||||
| 왜 | |
|
||||
|---|---|
|
||||
| keycloak-0 | kubelet 이 없어 **상태를 갱신할 수 없다.** 마지막으로 보고한 값이 그대로 얼어 있다 |
|
||||
| keycloak-1 | 살아서 **정직하게 보고한다** — DB 가 없으니 readiness 실패 |
|
||||
|
||||
> **파드 상태는 "지금 어떤가"가 아니라 "마지막으로 그렇게 들었다"이다.**
|
||||
> 노드가 죽으면 그 노드 파드의 상태는 **화석**이 된다.
|
||||
|
||||
Prometheus 는 정확했다.
|
||||
|
||||
```
|
||||
up{job=keycloak pod=keycloak-1} = 1
|
||||
up{job=keycloak pod=keycloak-0} = 0 ← 긁기에 실패했다 = 사실
|
||||
up{job=kubelet } = 0
|
||||
up{job=node-exporter node=kc-lab-2 } = 0
|
||||
```
|
||||
|
||||
**A-2 와 정반대다.** A-2 에서는 `up=1` 인데 서비스가 죽었고, 여기서는 `up=0`
|
||||
이 정확했다.
|
||||
|
||||
| | `up` 이 잡는가 |
|
||||
|---|---|
|
||||
| **대상이 사라짐** (노드 상실) | **잡는다** |
|
||||
| **대상이 살아서 못 씀** (DB 상실) | **못 잡는다** |
|
||||
|
||||

|
||||
|
||||
### 발견 ③ — 축출은 **5분** 뒤에 시작된다
|
||||
|
||||
쿠버네티스가 노드에 taint 를 붙인다.
|
||||
|
||||
```
|
||||
node.kubernetes.io/unreachable=:NoSchedule
|
||||
node.kubernetes.io/unreachable=:NoExecute
|
||||
```
|
||||
|
||||
그런데 모든 파드에는 기본 관용이 붙어 있다.
|
||||
|
||||
```
|
||||
node.kubernetes.io/not-ready NoExecute tolerationSeconds=300
|
||||
node.kubernetes.io/unreachable NoExecute tolerationSeconds=300
|
||||
```
|
||||
|
||||
```
|
||||
+240초 전부 Running
|
||||
+270초 a2-probe:Terminating keycloak-0:Terminating postgres:Terminating
|
||||
postgres-...-9cmsv:Pending ← 새 파드가 생겼다
|
||||
```
|
||||
|
||||
**노드가 죽고 약 5분이 지나야 파드가 축출된다.** 잠깐 끊긴 노드가 돌아올 수도
|
||||
있으니 성급하게 옮기지 않겠다는 설계다. 대신 **그동안은 아무 일도 일어나지
|
||||
않는다.**
|
||||
|
||||
### 발견 ④ — 축출돼도 **갈 곳이 없다**
|
||||
|
||||
```
|
||||
Warning FailedScheduling default-scheduler
|
||||
0/2 nodes are available:
|
||||
1 node(s) didn't match PersistentVolume's node affinity,
|
||||
1 node(s) had untolerated taint(s).
|
||||
```
|
||||
|
||||
```
|
||||
kc-lab-2 → taint 때문에 못 간다 (죽은 노드)
|
||||
kc-lab-1 → PVC 가 kc-lab-2 에 묶여 있어 못 간다
|
||||
```
|
||||
|
||||
**`local-path` PVC 는 그 노드의 로컬 디스크 디렉터리다.** 노드가 죽으면
|
||||
볼륨도 죽는다. 새 파드는 **영원히 Pending** 이다.
|
||||
|
||||
> 결함이 아니라 **조건**이다. 이 실험대는 그걸 알고 `local-path` 를 골랐다.
|
||||
> 운영이라면 네트워크 스토리지나 DB 복제가 이 자리를 메워야 한다.
|
||||
|
||||
### 발견 ⑤ — StatefulSet 은 **대체 파드를 만들지 않는다**
|
||||
|
||||
```
|
||||
NAME DESIRED READY CURRENT
|
||||
keycloak 2 <none> 1
|
||||
|
||||
keycloak-0 1/1 Terminating 30m ← 30분째
|
||||
keycloak-1 0/1 Running
|
||||
```
|
||||
|
||||
**`keycloak-0` 이 `Terminating` 에서 영원히 멈춰 있고, 대체 파드가 안 생긴다.**
|
||||
|
||||
| 왜 | |
|
||||
|---|---|
|
||||
| StatefulSet 의 계약 | **같은 이름의 파드는 클러스터에 하나뿐**이어야 한다 |
|
||||
| 컨트롤 플레인이 아는 것 | 노드가 안 보인다 = **파드가 죽었는지 확신할 수 없다** |
|
||||
| 그래서 | 옛 파드를 확실히 지우기 전엔 새 `keycloak-0` 을 못 만든다 |
|
||||
|
||||
**Deployment 였다면 즉시 새 파드를 만든다.** 이름이 아무래도 되기 때문이다.
|
||||
StatefulSet 의 "안정된 이름"이라는 이득의 **반대편 비용**이 여기다.
|
||||
|
||||
강제로 진행시키려면 사람이 개입해야 한다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab delete pod keycloak-0 --grace-period=0 --force
|
||||
```
|
||||
|
||||
**위험한 명령이다.** 노드가 사실은 살아 있고 네트워크만 끊긴 것이라면
|
||||
**같은 이름의 파드 둘이 동시에 존재**하게 된다 (split brain).
|
||||
|
||||
### 복구
|
||||
|
||||
```bash
|
||||
virsh start kc-lab-2
|
||||
```
|
||||
|
||||
```
|
||||
+30초 node=Ready | Running 파드 3 개 | 외부 HTTP 503
|
||||
+60초 node=Ready | Running 파드 3 개 | 외부 HTTP 200
|
||||
→ 서비스 복귀
|
||||
```
|
||||
|
||||
**60초.** 사람 개입 없이 전부 제자리로 돌아왔다.
|
||||
|
||||
---
|
||||
|
||||
## 4b. 컨트롤 플레인 노드 상실 (`kc-lab-1`)
|
||||
|
||||
### 이 노드에 무엇이 있었나 — 그게 곧 영향 범위다
|
||||
|
||||
```
|
||||
keycloak-lab keycloak-1
|
||||
kube-system coredns
|
||||
kube-system local-path-provisioner
|
||||
kube-system metrics-server
|
||||
kube-system svclb-traefik
|
||||
kube-system traefik ← replicas=1
|
||||
observability grafana
|
||||
observability prometheus
|
||||
observability node-exporter
|
||||
```
|
||||
|
||||
```
|
||||
NAME REPLICAS READY
|
||||
traefik 1 1 ← 진입점이 한 개다
|
||||
```
|
||||
|
||||
### 주입과 관찰
|
||||
|
||||
```bash
|
||||
virsh destroy kc-lab-1
|
||||
```
|
||||
|
||||
```
|
||||
+20초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server
|
||||
+60초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server
|
||||
+120초 외부 auth=000 grafana=502 | kubectl: Unable to connect to the server
|
||||
+160초 외부 auth=000 grafana=000 | kubectl: Unable to connect to the server
|
||||
```
|
||||
|
||||
### 발견 ⑥ — 워크로드는 살아 있는데 **길이 없다**
|
||||
|
||||
살아남은 노드에서 직접 확인했다.
|
||||
|
||||
```bash
|
||||
ssh kc-lab-2 'sudo crictl ps --name keycloak'
|
||||
```
|
||||
|
||||
```
|
||||
CONTAINER STATE NAME POD NAMESPACE
|
||||
e5f777900b76 Running keycloak keycloak-0 keycloak-lab
|
||||
```
|
||||
|
||||
**`keycloak-0` 은 멀쩡히 돌고 있다.** 컨테이너 런타임(containerd)은 API 서버
|
||||
없이도 이미 떠 있는 파드를 계속 굴린다.
|
||||
|
||||
```
|
||||
죽은 것: API 서버 · 스케줄러 · coredns · Traefik · Prometheus · Grafana
|
||||
산 것: keycloak-0 · postgres · containerd
|
||||
문제: 들어갈 문(Traefik)이 없다
|
||||
```
|
||||
|
||||
> **컨트롤 플레인 상실은 워크로드 상실이 아니다.**
|
||||
> 이미 떠 있는 것은 계속 돈다. **새로 뜨거나 옮기거나 고치는 것이 안 될 뿐.**
|
||||
|
||||
### 발견 ⑦ — `000` 과 `503` 의 차이
|
||||
|
||||
| 코드 | 뜻 |
|
||||
|---|---|
|
||||
| **`503`** (4a) | nginx·Traefik 은 살아 있고 **뒤에 보낼 파드가 없다** |
|
||||
| **`000`** (4b) | curl 이 응답 자체를 못 받았다 — **Traefik 이 없어 nginx 가 죽은 주소를 기다리다 타임아웃** |
|
||||
|
||||
`grafana=502` 가 한 번 찍힌 것이 증거다 — nginx 는 살아서 502 를 만들 수
|
||||
있었지만, 대부분은 타임아웃(`--max-time 8`)에 걸렸다.
|
||||
|
||||
**두 코드가 다른 층의 고장을 가리킨다.**
|
||||
|
||||
### 발견 ⑧ — 관측 시스템이 같이 죽으면 **0 이 아니라 구멍이 남는다**
|
||||
|
||||
Prometheus 가 `kc-lab-1` 에 있었다. 그 노드가 죽은 12:18–12:23 구간은
|
||||
그래프에 **0 이 아니라 데이터 없음**으로 남는다.
|
||||
|
||||
```
|
||||
대상이 죽음 → up = 0 → "언제 죽었는지" 알 수 있다
|
||||
관측자가 죽음 → 데이터 없음 → "그때 무슨 일이 있었는지" 모른다
|
||||
```
|
||||
|
||||
**계획 단계에서 "관측 스택은 죽이지 않는 노드에 둔다"고 정했던 이유**를
|
||||
실제로 확인한 셈이다. 다만 노드가 둘뿐이라 완전히 피할 수는 없다.
|
||||
|
||||
### 복구
|
||||
|
||||
```bash
|
||||
virsh start kc-lab-1
|
||||
```
|
||||
|
||||
```
|
||||
+30초 외부=502 | kc-lab-1=Ready kc-lab-2=Ready
|
||||
+60초 외부=200 | kc-lab-1=Ready kc-lab-2=Ready
|
||||
→ 서비스 복귀 (총 60초)
|
||||
```
|
||||
|
||||
**Grafana 는 로그인 세션을 잃었다** — 데이터가 `emptyDir` 이라 파드 재시작에
|
||||
사라진다. Prometheus 는 PVC 라 지표가 남았다. **의도한 설계대로 동작했다.**
|
||||
|
||||
---
|
||||
|
||||
## 5. 개념
|
||||
|
||||
### `node-monitor-grace-period` 와 `tolerationSeconds`
|
||||
|
||||
```
|
||||
기계 정지
|
||||
│
|
||||
│ 40초 node-monitor-grace-period → 노드 NotReady
|
||||
│
|
||||
│ +300초 tolerationSeconds (NoExecute) → 파드 축출 시작
|
||||
│
|
||||
▼
|
||||
총 약 5분 40초 동안 쿠버네티스는 아무것도 하지 않는다
|
||||
```
|
||||
|
||||
**빠른 장애 대응이 필요하면 이 값을 줄여야 하지만**, 짧게 하면 일시적
|
||||
네트워크 끊김에도 파드를 옮기게 되어 **불안정해진다.** 맞바꿈이다.
|
||||
|
||||
### `Terminating` 이 안 끝나는 이유
|
||||
|
||||
```
|
||||
파드 삭제 요청
|
||||
└─ kubelet 이 컨테이너를 멈추고 "지웠다"고 보고해야 끝난다
|
||||
└─ kubelet 이 없다 → 보고가 없다 → 영원히 Terminating
|
||||
```
|
||||
|
||||
`--grace-period=0 --force` 는 **보고를 기다리지 않고 API 에서 지우는 것**이다.
|
||||
컨테이너가 실제로 죽었는지는 **모른 채** 진행한다.
|
||||
|
||||
### StatefulSet vs Deployment — 노드 상실에서
|
||||
|
||||
| | Deployment | StatefulSet |
|
||||
|---|---|---|
|
||||
| 대체 파드 | **즉시 만든다** (이름 무관) | **안 만든다** (이름이 유일해야) |
|
||||
| 노드 상실 시 | 자동 회복 | **사람이 강제 삭제해야** |
|
||||
| 위험 | 없음 | 강제 삭제 시 **중복 실행 가능** |
|
||||
|
||||
**A-0 에서 "로그 대조가 쉬워서" StatefulSet 을 골랐는데, 그 대가가 여기 있다.**
|
||||
|
||||
### `local-path` PVC 와 재배치
|
||||
|
||||
```
|
||||
PVC → PV → nodeAffinity: kubernetes.io/hostname = kc-lab-2
|
||||
└─ 그 노드의 /var/lib/rancher/k3s/storage/... 디렉터리
|
||||
```
|
||||
|
||||
**볼륨이 노드에 못박히면 파드도 못박힌다.** 노드가 죽으면 데이터도 함께
|
||||
접근 불가가 된다.
|
||||
|
||||
---
|
||||
|
||||
## 6. 재현 절차 (명령어)
|
||||
|
||||
```bash
|
||||
export LIBVIRT_DEFAULT_URI=qemu:///system
|
||||
|
||||
# 사전 — PVC 가 어디 묶여 있는지 먼저 본다 (재배치 가능성이 여기서 정해진다)
|
||||
kubectl -n keycloak-lab get pvc -o name | while read p; do
|
||||
V=$(kubectl -n keycloak-lab get $p -o jsonpath='{.spec.volumeName}')
|
||||
kubectl get pv $V -o jsonpath='{.spec.nodeAffinity.required.nodeSelectorTerms[0].matchExpressions[0].values[0]}'
|
||||
done
|
||||
|
||||
# 4a 워커 상실
|
||||
virsh destroy kc-lab-2
|
||||
kubectl get node kc-lab-2 # 40초 뒤 NotReady
|
||||
kubectl -n keycloak-lab get pods -o wide # Running 인 채로 얼어 있다
|
||||
kubectl get node kc-lab-2 -o jsonpath='{.spec.taints}'
|
||||
# 5분 뒤 Terminating + 새 파드 Pending
|
||||
kubectl -n keycloak-lab describe pod <new-pod> | grep -A4 Events
|
||||
|
||||
# 4b 컨트롤 플레인 상실 — kubectl 이 죽으므로 노드에서 직접 본다
|
||||
virsh destroy kc-lab-1
|
||||
ssh kc-lab-2 'sudo crictl ps' # 워크로드는 살아 있다
|
||||
|
||||
# 복구
|
||||
virsh start kc-lab-2 && virsh start kc-lab-1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. 다음 실험에 남기는 것
|
||||
|
||||
| 실험 | 이 실험이 준 것 |
|
||||
|---|---|
|
||||
| **A-5** 비대칭 파티션 | 여기서는 노드가 **완전히** 사라졌다. 부분 단절은 더 고약하다 |
|
||||
| **D-1** 백업·복구 | **PVC 가 노드에 묶여 있다** — 노드가 영영 안 돌아오면 백업이 유일한 길 |
|
||||
| 구성 개선 | **Traefik `replicas=1` 은 진입점 단일 장애점** — 2로 늘리거나 DaemonSet 으로 |
|
||||
| 구성 개선 | **Grafana 가 `emptyDir`** — 재시작마다 로그인 세션이 사라진다 |
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,377 @@
|
||||
# Keycloak 멀티노드 클러스터 — 구성과 형성 확인
|
||||
|
||||
로드맵 1번. 세션 저장소 실험 전부의 선행 인프라다.
|
||||
브랜치 `feature/keycloak-multinode-cluster-jdbc-ping`.
|
||||
|
||||
**결과 — 두 파드가 서로 다른 노드에서 하나의 Infinispan 클러스터를 이뤘다.**
|
||||
|
||||
---
|
||||
|
||||
## 1. 무엇을 확인하려는가
|
||||
|
||||
Keycloak 26은 **디스커버리와 클러스터 통신을 서로 다른 경로로** 처리한다.
|
||||
|
||||
| 단계 | 경로 | 실패하면 |
|
||||
|---|---|---|
|
||||
| **디스커버리** — 서로를 찾는다 | PostgreSQL의 `JGROUPS_PING` 테이블 | 상대의 존재 자체를 모른다 |
|
||||
| **클러스터 통신** — 실제로 대화한다 | **TCP 7800** (파드 간 직접) | **DB에는 등록되는데 클러스터가 안 붙는다** |
|
||||
|
||||
두 번째 줄이 이 실험대를 2노드로 만든 이유다. **단일 노드에서는 이 고장을
|
||||
재현할 수 없다** — 같은 커널 안에서는 막을 경계가 없기 때문이다.
|
||||
|
||||
먼저 **정상적으로 붙는 상태**를 확보하고 실측값을 남긴다. 그래야 다음 실험에서
|
||||
깨뜨렸을 때 무엇이 달라졌는지 비교할 수 있다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 배포한 구성과 그 근거
|
||||
|
||||
매니페스트: [`deploy/lab/k8s/keycloak-cluster.yaml`](../deploy/lab/k8s/keycloak-cluster.yaml)
|
||||
|
||||
### 2-1. 왜 StatefulSet인가
|
||||
|
||||
Deployment를 쓰면 파드 이름이 `keycloak-7d9f8b-x4k2p`처럼 매번 바뀐다.
|
||||
StatefulSet은 **`keycloak-0`, `keycloak-1`로 고정**된다.
|
||||
|
||||
```yaml
|
||||
kind: StatefulSet
|
||||
spec:
|
||||
serviceName: keycloak-headless
|
||||
replicas: 2
|
||||
podManagementPolicy: Parallel
|
||||
```
|
||||
|
||||
**이 실험에서 이름 안정성이 중요한 이유** — 클러스터 멤버십을 읽는 곳이 두
|
||||
군데인데(Infinispan 로그, `JGROUPS_PING` 테이블) 이름이 계속 바뀌면 대조가
|
||||
어렵다. 실제로 Infinispan은 `keycloak-0-49501`처럼 **파드 이름 + 랜덤 접미사**를
|
||||
노드 식별자로 쓴다.
|
||||
|
||||
**`podManagementPolicy: Parallel`** — 기본값 `OrderedReady`는 0번이 Ready가 된
|
||||
뒤에야 1번을 만든다. `Parallel`은 **동시에 시작**하므로 두 파드가 DB에 등록을
|
||||
경쟁하게 되고, 그것이 운영에서 실제로 일어나는 상황이다.
|
||||
|
||||
### 2-2. 왜 `start`이고 `start-dev`가 아닌가
|
||||
|
||||
```yaml
|
||||
args: ["start"]
|
||||
```
|
||||
|
||||
`start-dev`는 **`cache=local`을 강제**한다. 클러스터가 아예 형성되지 않는다.
|
||||
저장소의 `docker-compose.yml`이 `start-dev`를 쓰는 것은 단일 인스턴스 학습용이며,
|
||||
이 실험대에서는 쓸 수 없다.
|
||||
|
||||
`--optimized`는 붙이지 않았다. 붙이려면 사전 `build`가 필요하고, 없으면
|
||||
첫 기동에 **암묵적 build가 실행되어 60~90초**가 걸린다. 그래서 아래처럼
|
||||
`startupProbe`를 넉넉하게 준다.
|
||||
|
||||
### 2-3. 노드당 하나씩 배치
|
||||
|
||||
```yaml
|
||||
topologySpreadConstraints:
|
||||
- maxSkew: 1
|
||||
topologyKey: kubernetes.io/hostname
|
||||
whenUnsatisfiable: ScheduleAnyway
|
||||
labelSelector:
|
||||
matchLabels: { app: keycloak }
|
||||
```
|
||||
|
||||
**두 파드가 한 노드에 몰리면 7800 차단 실험이 무의미해진다.** 같은 커널 안의
|
||||
루프백 통신이라 막을 대상이 없기 때문이다.
|
||||
|
||||
`ScheduleAnyway`를 고른 이유는 장애 실험 때문이다. `DoNotSchedule`이면 노드
|
||||
하나를 죽였을 때 남은 파드가 **배치되지 못하고 Pending에 머문다.**
|
||||
|
||||
### 2-4. 헬스체크는 9000 포트다
|
||||
|
||||
```yaml
|
||||
ports:
|
||||
- { containerPort: 8080, name: http }
|
||||
- { containerPort: 9000, name: management }
|
||||
- { containerPort: 7800, name: jgroups }
|
||||
|
||||
startupProbe: { httpGet: { path: /health/started, port: management }, failureThreshold: 60 }
|
||||
readinessProbe:{ httpGet: { path: /health/ready, port: management } }
|
||||
livenessProbe: { httpGet: { path: /health/live, port: management } }
|
||||
```
|
||||
|
||||
**Keycloak 25부터 health와 metrics가 8080이 아니라 관리 포트 9000으로 옮겨졌다.**
|
||||
8080으로 프로브를 걸면 404가 나고 파드가 영원히 Ready가 되지 않는다.
|
||||
|
||||
`KC_HEALTH_ENABLED=true`를 켜야 엔드포인트가 노출된다.
|
||||
|
||||
`startupProbe`의 `failureThreshold: 60` × `periodSeconds: 10` = **최대 10분**을
|
||||
기다린다. 첫 기동의 암묵적 build 때문이다. 이게 없으면 liveness가 먼저 발동해
|
||||
**재시작 루프**에 빠진다.
|
||||
|
||||
### 2-5. 환경변수 — 첫 실험에서 확정한 값
|
||||
|
||||
```yaml
|
||||
- { name: KC_HOSTNAME, value: https://auth.hyeonworks.com }
|
||||
- { name: KC_HOSTNAME_STRICT, value: "true" }
|
||||
- { name: KC_PROXY_HEADERS, value: xforwarded }
|
||||
- { name: KC_HTTP_ENABLED, value: "true" }
|
||||
```
|
||||
|
||||
[`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md)에서
|
||||
측정으로 확정한 조합이다.
|
||||
|
||||
| 설정 | 역할 |
|
||||
|---|---|
|
||||
| `KC_HOSTNAME`에 **전체 URL** | 스킴·호스트를 **고정**한다. 헤더와 무관하게 `iss`가 https로 발급된다 |
|
||||
| `KC_HOSTNAME_STRICT=true` | Host 헤더를 믿지 않는다. 조작으로 흐름을 돌릴 여지를 없앤다 |
|
||||
| `KC_PROXY_HEADERS=xforwarded` | **클라이언트 IP** 등 나머지를 forwarded 헤더에서 가져온다 |
|
||||
| `KC_HTTP_ENABLED=true` | 앞단이 TLS를 끊었으므로 평문 HTTP를 받는다 |
|
||||
|
||||
**이 실험을 먼저 하지 않았다면** 지금 `iss`가 `http://10.42.x.x`로 나왔을 것이고,
|
||||
원인을 세션 쪽에서 찾느라 헤맸을 것이다.
|
||||
|
||||
### 2-6. 힙 상한
|
||||
|
||||
```yaml
|
||||
- { name: JAVA_OPTS_KC_HEAP, value: "-Xms256m -Xmx512m" }
|
||||
resources:
|
||||
requests: { memory: 640Mi, cpu: 100m }
|
||||
limits: { memory: 900Mi }
|
||||
```
|
||||
|
||||
Keycloak은 기본값이 넉넉해 그냥 두면 1GB를 넘긴다. 이 실험대의 게스트 여유가
|
||||
약 3.8GB이므로 명시적으로 잡는다. 실측 결과 **파드당 약 590Mi**로 안정됐다.
|
||||
|
||||
### 2-7. PostgreSQL — 볼륨이 노드에 고정된다
|
||||
|
||||
```yaml
|
||||
storageClassName: local-path
|
||||
strategy:
|
||||
type: Recreate
|
||||
env:
|
||||
- { name: PGDATA, value: /var/lib/postgresql/data/pgdata }
|
||||
```
|
||||
|
||||
k3s 기본 `local-path` 프로비저너는 **파드가 배치된 노드의 로컬 디스크**에
|
||||
볼륨을 만든다. 따라서 PostgreSQL은 그 노드에 묶인다.
|
||||
|
||||
**이것은 결함이 아니라 실험 조건이다.** 나중에 "데이터베이스가 있는 노드가
|
||||
죽으면" 시나리오가 그래서 의미를 갖는다.
|
||||
|
||||
- `strategy: Recreate` — RWO 볼륨은 두 파드가 동시에 마운트할 수 없다.
|
||||
기본값 `RollingUpdate`면 새 파드가 볼륨을 못 잡고 멈춘다
|
||||
- `PGDATA`를 한 단계 아래로 — 마운트 지점에 `lost+found` 같은 것이 있으면
|
||||
`initdb`가 거부한다
|
||||
|
||||
### 2-8. 헤드리스 서비스는 왜 두는가
|
||||
|
||||
```yaml
|
||||
kind: Service
|
||||
metadata: { name: keycloak-headless }
|
||||
spec:
|
||||
clusterIP: None
|
||||
```
|
||||
|
||||
**jdbc-ping 디스커버리에는 필요 없다.** DB로 서로를 찾기 때문이다.
|
||||
개별 파드에 안정된 DNS 이름으로 접근해 상태를 조회하기 위해 둔다.
|
||||
|
||||
---
|
||||
|
||||
## 3. 실행한 명령
|
||||
|
||||
### 3-1. 브랜치와 정리
|
||||
|
||||
```bash
|
||||
# 워크스테이션
|
||||
cd ~/workspace/keycloak-pattern
|
||||
git checkout -b feature/keycloak-multinode-cluster-jdbc-ping
|
||||
git merge --no-edit develop-keycloak-session-store
|
||||
|
||||
# lab host — 끝난 실험을 지워 메모리를 회수한다
|
||||
kubectl delete ns header-lab
|
||||
```
|
||||
|
||||
정리 후 게스트 사용량이 `kc-lab-1 1593Mi(46%)` / `kc-lab-2 872Mi(35%)`로 떨어졌다.
|
||||
|
||||
### 3-2. 배포
|
||||
|
||||
```bash
|
||||
# 워크스테이션 — 매니페스트 작성 후
|
||||
git add deploy/lab/k8s/keycloak-cluster.yaml
|
||||
git commit -m "feat: deploy Keycloak multi-node cluster with PostgreSQL"
|
||||
git push -u origin feature/keycloak-multinode-cluster-jdbc-ping
|
||||
|
||||
# lab host
|
||||
cd ~/workspace/keycloak-pattern
|
||||
git fetch origin
|
||||
git checkout -b feature/keycloak-multinode-cluster-jdbc-ping origin/feature/keycloak-multinode-cluster-jdbc-ping
|
||||
kubectl apply -f deploy/lab/k8s/keycloak-cluster.yaml
|
||||
```
|
||||
|
||||
**PostgreSQL을 먼저 기다린다.** Keycloak이 DB 없이 뜨면 기동에 실패한다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab rollout status deployment/postgres --timeout=180s
|
||||
kubectl -n keycloak-lab rollout status statefulset/keycloak --timeout=600s
|
||||
```
|
||||
|
||||
이미지를 당겨오고 암묵적 build가 도는 첫 기동은 **수 분** 걸린다.
|
||||
|
||||
### 3-3. 검증
|
||||
|
||||
```bash
|
||||
# 파드 배치 — 서로 다른 노드에 있어야 한다
|
||||
kubectl -n keycloak-lab get pods -o wide
|
||||
|
||||
# 클러스터 뷰 — Infinispan 로그
|
||||
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
|
||||
|
||||
# 디스커버리 테이블
|
||||
PG=$(kubectl -n keycloak-lab get pod -l app=postgres -o name | head -1)
|
||||
kubectl -n keycloak-lab exec "$PG" -- \
|
||||
psql -U keycloak -d keycloak -c "SELECT name, cluster_name, ip, coord FROM jgroups_ping ORDER BY name;"
|
||||
|
||||
# 외부 접근과 issuer
|
||||
curl -s https://auth.hyeonworks.com/realms/master/.well-known/openid-configuration | python3 -m json.tool
|
||||
|
||||
# 자원
|
||||
kubectl -n keycloak-lab top pods
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 확인된 사실
|
||||
|
||||
증거 원자료: [`evidence/keycloak-multinode-cluster/`](evidence/keycloak-multinode-cluster/)
|
||||
|
||||
### 4-1. 클러스터가 형성됐다
|
||||
|
||||
```
|
||||
ISPN000094: Received new cluster view for channel ISPN:
|
||||
[keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938, keycloak-0-49501]
|
||||
↑ 멤버 수
|
||||
ISPN100000: Node keycloak-0-49501 joined the cluster
|
||||
ISPN000079: Channel `ISPN` local address is `keycloak-0-49501`,
|
||||
physical addresses are `[10.42.1.18:7800]`
|
||||
```
|
||||
|
||||
두 파드가 **동일한 뷰**를 보고 있고, 물리 주소가 **7800**임이 로그에 찍힌다.
|
||||
|
||||
### 4-2. 디스커버리와 통신이 분리되어 있다
|
||||
|
||||
```
|
||||
name | cluster_name | ip | coord
|
||||
------------------+--------------+-----------------+-------
|
||||
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
|
||||
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
|
||||
```
|
||||
|
||||
**테이블 하나에 두 메커니즘이 다 보인다.**
|
||||
|
||||
- `name`·`cluster_name` — **DB로 하는 디스커버리**의 결과
|
||||
- `ip` 컬럼의 `:7800` — **실제 통신이 일어날 경로**
|
||||
|
||||
`coord`가 `t`인 `keycloak-1`이 코디네이터다. 이 노드를 죽였을 때 인계가
|
||||
일어나는지가 다음 실험 항목이다.
|
||||
|
||||
전체 스키마는 `address / name / cluster_name / ip / coord / last_update /
|
||||
coordinated_by`이며 기본키는 `address`다.
|
||||
|
||||
### 4-3. 노드당 하나씩 배치됐다
|
||||
|
||||
```
|
||||
keycloak-0 10.42.1.18 kc-lab-2
|
||||
keycloak-1 10.42.0.16 kc-lab-1
|
||||
postgres 10.42.1.19 kc-lab-2
|
||||
```
|
||||
|
||||
파드 IP 대역이 노드를 알려준다(`10.42.0.x` = kc-lab-1, `10.42.1.x` = kc-lab-2).
|
||||
**독립된 커널 두 개에 하나씩** 떴으므로 7800 차단 실험의 전제가 성립한다.
|
||||
|
||||
PostgreSQL이 `kc-lab-2`에 있다는 점도 기록해둔다. **`kc-lab-2`를 죽이면
|
||||
Keycloak 하나와 데이터베이스가 동시에 사라진다.**
|
||||
|
||||
### 4-4. 2홉 헤더 계약이 실제로 작동한다
|
||||
|
||||
```
|
||||
issuer https://auth.hyeonworks.com/realms/master
|
||||
authorization_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/auth
|
||||
token_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/token
|
||||
end_session_endpoint https://auth.hyeonworks.com/realms/master/protocol/openid-connect/logout
|
||||
jwks_uri https://auth.hyeonworks.com/realms/master/protocol/openid-connect/certs
|
||||
```
|
||||
|
||||
**전부 `https`이고 외부 호스트명이다.** 첫 실험의 결론이 그대로 값을 했다.
|
||||
|
||||
### 4-5. 자원
|
||||
|
||||
```
|
||||
keycloak-0 594Mi
|
||||
keycloak-1 593Mi
|
||||
postgres 67Mi
|
||||
──────────────────────
|
||||
kc-lab-1 2248Mi (65%)
|
||||
kc-lab-2 1447Mi (58%)
|
||||
```
|
||||
|
||||
예상(파드당 700Mi)보다 적다. `JAVA_OPTS_KC_HEAP` 제한이 작동했다.
|
||||
BFF와 Redis를 추가할 여유가 남아 있다.
|
||||
|
||||
---
|
||||
|
||||
## 5. 겪은 함정
|
||||
|
||||
### `JGROUPS_PING` 컬럼명은 자료마다 다르다
|
||||
|
||||
오래된 문서에는 `own_addr`, `ping_data` 같은 이름이 나오지만 **Keycloak 26의
|
||||
실제 스키마는 다르다.**
|
||||
|
||||
```
|
||||
address / name / cluster_name / ip / coord / last_update / coordinated_by
|
||||
```
|
||||
|
||||
쿼리 전에 `\d jgroups_ping`으로 확인한다.
|
||||
|
||||
### Keycloak 컨테이너에 `curl`이 없다
|
||||
|
||||
메트릭을 파드 안에서 조회하려다 실패했다.
|
||||
|
||||
```
|
||||
sh: line 1: curl: command not found
|
||||
```
|
||||
|
||||
Keycloak 공식 이미지는 최소 구성이다. 메트릭을 볼 때는 포트포워딩하거나
|
||||
임시 파드를 쓴다.
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab port-forward keycloak-0 9000:9000 &
|
||||
curl -s localhost:9000/metrics | grep -i cluster
|
||||
|
||||
# 또는
|
||||
kubectl -n keycloak-lab run m --rm -i --restart=Never --image=curlimages/curl:8.11.1 -- \
|
||||
curl -s http://keycloak-0.keycloak-headless:9000/metrics
|
||||
```
|
||||
|
||||
### 첫 기동이 느린 것은 정상이다
|
||||
|
||||
`--optimized` 없이 `start`하면 **암묵적 build**가 실행된다. `startupProbe`를
|
||||
넉넉히 주지 않으면 liveness가 먼저 발동해 재시작 루프에 빠진다.
|
||||
|
||||
---
|
||||
|
||||
## 6. 다음 실험 — 깨뜨려서 무엇이 보이는지
|
||||
|
||||
정상 상태를 확보했으므로 이제 의도적으로 고장을 만든다.
|
||||
|
||||
| 실험 | 방법 | 확인할 것 |
|
||||
|---|---|---|
|
||||
| **7800 차단** | NetworkPolicy로 파드 간 7800만 차단 | **DB엔 등록되는데 클러스터가 안 붙는** 증상. 로그에 무엇이 먼저 보이는가 |
|
||||
| **노드 상실** | `virsh destroy kc-lab-2` | 코디네이터 인계가 일어나는가. PostgreSQL도 같이 죽는다는 점에 유의 |
|
||||
| **DB 상실** | postgres 파드 정지 | 이미 형성된 클러스터는 버티는가. 새 로그인은? |
|
||||
|
||||
**7800 차단부터 하는 것이 좋다.** 되돌리기가 가장 쉽고(NetworkPolicy 삭제),
|
||||
증상이 로그에 선명하게 남는다.
|
||||
|
||||
## 참고
|
||||
|
||||
| 문서 | 관계 |
|
||||
|---|---|
|
||||
| [`session-store-lab-roadmap.md`](session-store-lab-roadmap.md) | 이 실험은 로드맵 1번 |
|
||||
| [`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md) | `KC_HOSTNAME`·`KC_PROXY_HEADERS` 값의 근거 |
|
||||
| [`session-lab-operations.md`](session-lab-operations.md) | 명령·자원 예산 |
|
||||
| [`session-lab-concepts.md`](session-lab-concepts.md) | StatefulSet·프로브·PVC 등 개념 |
|
||||
@@ -0,0 +1,429 @@
|
||||
# 관측성 — Prometheus · node-exporter · Grafana
|
||||
|
||||
로드맵 10번. 장애 주입 실험보다 **먼저** 세운다.
|
||||
|
||||
**왜 먼저인가** — 나중에 세우면 이미 지나간 장애의 지표를 볼 수 없다.
|
||||
"클러스터가 1분쯤 뒤에 복구됐다"는 측정이 아니라 인상이다.
|
||||
로드맵에 *"장애 주입 중에 어떤 지표가 먼저 움직이는지 기록한다"*고 적어둔 항목은
|
||||
관측이 먼저 서 있어야만 가능하다.
|
||||
|
||||
메모리를 8GB → 12GB로 증설한 뒤에야 올릴 수 있게 됐다.
|
||||
|
||||
---
|
||||
|
||||
## 1. 무엇을 세웠나
|
||||
|
||||
```
|
||||
┌─ Grafana ──────────┐
|
||||
브라우저 ──────▶│ app2.hyeonworks.com│ 대시보드
|
||||
└─────────┬──────────┘
|
||||
│ PromQL
|
||||
┌─────────▼──────────┐
|
||||
│ Prometheus │ 수집·저장 (TSDB, 7일)
|
||||
└─────────┬──────────┘
|
||||
│ scrape (15초)
|
||||
┌───────────────────┼───────────────────┐
|
||||
▼ ▼ ▼
|
||||
Keycloak :9000 node-exporter :9100 kubelet
|
||||
(앱 지표) (머신 지표) (컨테이너 지표)
|
||||
```
|
||||
|
||||
매니페스트: [`deploy/lab/k8s/observability.yaml`](../deploy/lab/k8s/observability.yaml)
|
||||
|
||||
| 구성요소 | 역할 | 실측 메모리 |
|
||||
|---|---|---|
|
||||
| Prometheus | 수집·저장·질의 | 164Mi |
|
||||
| node-exporter (DaemonSet) | 노드당 하나, 머신 지표 | 8Mi × 2 |
|
||||
| Grafana | 시각화 | 65Mi |
|
||||
| **합계** | | **약 245Mi** |
|
||||
|
||||
예상(550Mi)보다 훨씬 적다. 실험대 규모에서는 관측성 비용이 거의 무시할 수준이다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 왜 kube-prometheus-stack을 쓰지 않았나
|
||||
|
||||
Helm 차트 하나로 끝내는 방법이 있지만 **평범한 매니페스트를 직접 썼다.**
|
||||
|
||||
| | kube-prometheus-stack | 직접 작성 |
|
||||
|---|---|---|
|
||||
| 설치 | Helm 한 줄 | 매니페스트 400줄 |
|
||||
| 메모리 | 1.5GB 이상 | **245Mi** |
|
||||
| 포함 | Operator, Alertmanager, 대시보드 다수, kube-state-metrics | 필요한 것만 |
|
||||
| **보이는 것** | 추상화 뒤에 숨음 | **스크레이프 설정·RBAC·relabel 이 눈에 보임** |
|
||||
|
||||
세 번째 줄이 결정적이다. 이 실험대의 목적은 **인과를 직접 확인하는 것**이므로,
|
||||
"어떻게 타깃을 찾는가"가 YAML에 드러나 있어야 한다. Operator를 쓰면
|
||||
`ServiceMonitor` 하나만 보이고 그 아래는 감춰진다.
|
||||
|
||||
---
|
||||
|
||||
## 3. 구성 결정과 근거
|
||||
|
||||
### 3-1. 관측 스택의 배치 — 장애 도메인 분리
|
||||
|
||||
```yaml
|
||||
nodeSelector:
|
||||
node-role.kubernetes.io/control-plane: "true"
|
||||
```
|
||||
|
||||
**관측 시스템은 관측 대상과 같은 장애 도메인에 있으면 안 된다.** 죽는 순간을
|
||||
기록해야 하는데 같이 죽으면 기록이 남지 않는다.
|
||||
|
||||
노드가 둘뿐이라 완전히 피할 수는 없다. 그래서 규칙을 정했다.
|
||||
|
||||
| 노드 | 역할 | 실험에서 |
|
||||
|---|---|---|
|
||||
| **kc-lab-1** (k3s **server**) | control plane · Traefik · coredns · metrics-server · local-path-provisioner | **관측 스택을 여기 둔다. 죽이지 않는다** |
|
||||
| **kc-lab-2** (k3s **agent**) | keycloak-0 · postgres | **장애 주입 대상** |
|
||||
|
||||
`kubernetes.io/hostname`으로 못박지 않고 **`node-role.kubernetes.io/control-plane`
|
||||
라벨**을 쓴 이유는 의미가 드러나기 때문이다 — "컨트롤 플레인 노드에 둔다"는
|
||||
의도가 호스트 이름보다 오래간다.
|
||||
|
||||
### 3-2. 앞선 판단을 정정했다
|
||||
|
||||
배치를 조사하기 전에는 **"노드 상실 실험은 `kc-lab-1`을 죽여서 하자"**고
|
||||
적었다. 그 노드에 Keycloak 하나만 있다고 생각했기 때문이다. **틀렸다.**
|
||||
|
||||
```
|
||||
kc-lab-1 (server) keycloak-1, traefik, coredns, metrics-server, local-path-provisioner
|
||||
kc-lab-2 (agent) keycloak-0, postgres
|
||||
```
|
||||
|
||||
`kc-lab-1`을 죽이면 **API 서버·DNS·인그레스가 한꺼번에 사라진다.** 노드 상실이
|
||||
아니라 **컨트롤 플레인 상실**이며, `kubectl`조차 동작하지 않는다.
|
||||
|
||||
**깨끗한 워커 노드 상실 실험은 `kc-lab-2`를 죽이는 것이다.** 그때도 변수가
|
||||
둘(keycloak-0 + postgres)이지만, 클러스터 제어는 살아 있고 관측도 계속된다.
|
||||
|
||||
### 3-3. 스크레이프 주기 15초
|
||||
|
||||
```yaml
|
||||
global:
|
||||
scrape_interval: 15s
|
||||
```
|
||||
|
||||
운영에서는 30~60초가 흔하지만 여기서는 짧게 잡았다. **노드가 죽는 순간을
|
||||
두어 샘플 안에 잡아야** "무엇이 먼저 움직였나"를 말할 수 있다.
|
||||
60초면 장애와 복구가 같은 샘플에 뭉개진다.
|
||||
|
||||
### 3-4. 타깃을 정적 목록으로 두지 않는다
|
||||
|
||||
```yaml
|
||||
kubernetes_sd_configs:
|
||||
- role: endpoints
|
||||
namespaces: { names: [keycloak-lab] }
|
||||
```
|
||||
|
||||
**파드 IP는 재시작마다 바뀐다.** 실험대를 전원 종료했다 켰을 때 모든 파드가
|
||||
새 주소를 받는 것을 직접 확인했다(`10.42.1.22` → `10.42.1.25`).
|
||||
정적 목록을 적어두면 그때마다 깨진다.
|
||||
|
||||
쿠버네티스 API에 물어보는 방식(service discovery)이므로 **파드가 옮겨다녀도
|
||||
따라간다.** Traefik의 `trustedIPs`에 개별 IP를 적을 수 없었던 것과 같은 이유다.
|
||||
|
||||
### 3-5. relabel — 발견한 것을 걸러내고 이름을 붙인다
|
||||
|
||||
```yaml
|
||||
relabel_configs:
|
||||
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
|
||||
action: keep
|
||||
regex: keycloak-headless;management
|
||||
- source_labels: [__meta_kubernetes_pod_name]
|
||||
target_label: pod
|
||||
- source_labels: [__meta_kubernetes_pod_node_name]
|
||||
target_label: node
|
||||
```
|
||||
|
||||
service discovery는 네임스페이스의 **모든 엔드포인트**를 가져온다. 그중
|
||||
필요한 것만 남기고 나머지는 버리는 것이 `keep`이다.
|
||||
|
||||
- 첫 규칙 — `keycloak-headless` 서비스의 `management` 포트만 남긴다.
|
||||
8080(http)까지 긁으면 애플리케이션 트래픽 포트에 헛되이 요청이 간다
|
||||
- 나머지 두 규칙 — **`pod`과 `node` 라벨을 붙인다.** 이것이 없으면
|
||||
"어느 파드가, 어느 노드에서" 라는 질문에 답할 수 없다.
|
||||
노드 상실 실험에서 결정적이다
|
||||
|
||||
### 3-6. Keycloak 지표는 9000 포트다
|
||||
|
||||
헬스체크와 같은 관리 포트다. `KC_METRICS_ENABLED=true`가 이미 StatefulSet에
|
||||
설정돼 있다. **8080을 긁으면 지표가 나오지 않는다.**
|
||||
|
||||
### 3-7. node-exporter는 DaemonSet + 호스트 네임스페이스
|
||||
|
||||
```yaml
|
||||
kind: DaemonSet
|
||||
spec:
|
||||
template:
|
||||
spec:
|
||||
hostNetwork: true
|
||||
hostPID: true
|
||||
tolerations:
|
||||
- operator: Exists
|
||||
```
|
||||
|
||||
- **DaemonSet** — 노드마다 정확히 하나. 죽을 노드에도 있어야 **꺼지기 직전의
|
||||
마지막 샘플**이 남는다
|
||||
- **`hostNetwork`/`hostPID`** — 측정 대상이 컨테이너가 아니라 **머신**이다.
|
||||
컨테이너 네임스페이스 안에서 보면 자기 자신만 보인다
|
||||
- **`tolerations: operator: Exists`** — 어떤 taint가 걸린 노드에도 뜬다.
|
||||
관측이 빠지는 노드가 있으면 안 된다
|
||||
|
||||
### 3-8. Prometheus 저장소는 PVC
|
||||
|
||||
```yaml
|
||||
storageClassName: local-path
|
||||
--storage.tsdb.retention.time=7d
|
||||
```
|
||||
|
||||
`emptyDir`로 두면 파드가 재시작될 때 **장애 실험의 기록이 통째로 사라진다.**
|
||||
사후 추적이 목적이므로 영속 저장이 필요하다.
|
||||
|
||||
`local-path`는 노드에 고정되므로 Prometheus도 `kc-lab-1`에 묶인다.
|
||||
`nodeSelector`와 방향이 같아 문제가 되지 않는다.
|
||||
|
||||
보존 7일은 실험 기간보다 넉넉하면서 **볼륨이 노드를 채우는 원인이 되지 않을**
|
||||
크기다.
|
||||
|
||||
```yaml
|
||||
securityContext:
|
||||
fsGroup: 65534
|
||||
```
|
||||
|
||||
`prom/prometheus` 이미지는 `nobody`(65534)로 실행된다. `fsGroup`이 없으면
|
||||
새로 만들어진 볼륨의 소유자가 root라 **쓰기 권한이 없어 기동에 실패한다.**
|
||||
|
||||
### 3-9. Grafana에도 외부 URL을 알려줘야 한다
|
||||
|
||||
```yaml
|
||||
- name: GF_SERVER_ROOT_URL
|
||||
value: https://app2.hyeonworks.com
|
||||
```
|
||||
|
||||
**Keycloak의 `KC_HOSTNAME`과 정확히 같은 성격의 설정이다.** Grafana도
|
||||
리다이렉트와 자산 경로에 절대 URL을 만든다. 이 값이 없으면 로그인 리다이렉트가
|
||||
`http://<파드IP>:3000`으로 나간다.
|
||||
|
||||
2홉 헤더 계약에서 확인한 원리가 여기서도 그대로 적용된다 —
|
||||
**프록시 뒤의 애플리케이션은 자기가 외부에서 어떤 주소로 보이는지 모른다.**
|
||||
|
||||
### 3-10. 데이터소스는 파일로 프로비저닝
|
||||
|
||||
```yaml
|
||||
volumeMounts:
|
||||
- name: datasources
|
||||
mountPath: /etc/grafana/provisioning/datasources
|
||||
```
|
||||
|
||||
UI에서 클릭으로 추가하면 Grafana 자체 DB에만 남는다. 그 DB는 여기서
|
||||
`emptyDir`이므로 **파드가 재시작되면 사라진다.** 파일로 두면 항상 같은 상태로
|
||||
뜬다.
|
||||
|
||||
### 3-11. Grafana를 `app2`에 붙인 이유
|
||||
|
||||
인증서에 들어 있는 이름이 `auth` / `app1` / `app2` 셋뿐이고 `app2`가 비어
|
||||
있었다. **SSO 실험에서 `app2`가 필요해지면 옮긴다.**
|
||||
|
||||
---
|
||||
|
||||
## 4. 실행한 명령
|
||||
|
||||
```bash
|
||||
# 워크스테이션 — 매니페스트 작성 후
|
||||
git add deploy/lab/k8s/observability.yaml
|
||||
git commit -m "feat: add Prometheus, node-exporter and Grafana"
|
||||
git push origin feature/keycloak-multinode-cluster-jdbc-ping
|
||||
|
||||
# lab host
|
||||
cd ~/workspace/keycloak-pattern && git pull
|
||||
kubectl apply -f deploy/lab/k8s/observability.yaml
|
||||
|
||||
kubectl -n observability rollout status deployment/prometheus --timeout=300s
|
||||
kubectl -n observability rollout status daemonset/node-exporter --timeout=180s
|
||||
kubectl -n observability rollout status deployment/grafana --timeout=300s
|
||||
```
|
||||
|
||||
**검증 — 배포 성공과 타깃 수집은 다른 문제다.**
|
||||
|
||||
```bash
|
||||
kubectl -n observability run q --rm -i --restart=Never \
|
||||
--image=curlimages/curl:8.11.1 --quiet --command -- \
|
||||
curl -s "http://prometheus.observability.svc:9090/api/v1/targets?state=any" > /tmp/targets.json
|
||||
|
||||
python3 -c "
|
||||
import json
|
||||
d,_ = json.JSONDecoder().raw_decode(open('/tmp/targets.json').read())
|
||||
ts = d['data']['activeTargets']
|
||||
print(f\"{sum(1 for t in ts if t['health']=='up')}/{len(ts)} up\")
|
||||
for t in ts:
|
||||
if t['health'] != 'up': print(t['labels'], t.get('lastError'))
|
||||
"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. 겪은 함정
|
||||
|
||||
### kubelet 타깃이 403 Forbidden
|
||||
|
||||
첫 배포에서 **7개 중 5개만 up**이었다.
|
||||
|
||||
```
|
||||
DOWN kubelet kc-lab-1 server returned HTTP status 403 Forbidden
|
||||
DOWN kubelet kc-lab-2 server returned HTTP status 403 Forbidden
|
||||
```
|
||||
|
||||
원인은 RBAC였다. kubelet 지표는 **API 서버의 proxy 서브리소스**를 통해
|
||||
가져온다.
|
||||
|
||||
```
|
||||
/api/v1/nodes/<name>/proxy/metrics
|
||||
─────
|
||||
```
|
||||
|
||||
이 경로에는 `nodes`나 `nodes/metrics`가 아니라 **`nodes/proxy`** 권한이
|
||||
필요하다.
|
||||
|
||||
```diff
|
||||
- resources: [nodes, nodes/metrics, services, endpoints, pods]
|
||||
+ resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
|
||||
```
|
||||
|
||||
**다른 잡은 전부 정상이었다.** 이런 부분 실패는 타깃 목록을 직접 확인하지
|
||||
않으면 드러나지 않는다. `rollout status`는 "성공"이라고 말한다.
|
||||
|
||||
### `kubectl run --rm -i`의 출력에 종료 메시지가 섞인다
|
||||
|
||||
```
|
||||
json.decoder.JSONDecodeError: Extra data: line 1 column 54973
|
||||
```
|
||||
|
||||
`kubectl run --rm`은 컨테이너 출력 뒤에 `pod "q" deleted`를 덧붙인다.
|
||||
JSON 파서가 그 뒤를 만나면 실패한다.
|
||||
|
||||
**해결** — `raw_decode`로 앞쪽의 완전한 JSON만 읽는다.
|
||||
|
||||
```python
|
||||
d, _ = json.JSONDecoder().raw_decode(raw)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. 실험에 쓸 지표
|
||||
|
||||
메트릭 이름이 **1506개** 수집된다. 그중 장애 실험에서 볼 것들이다.
|
||||
|
||||
### 가장 중요한 것 — `up`
|
||||
|
||||
```promql
|
||||
up
|
||||
up{job="keycloak"}
|
||||
```
|
||||
|
||||
Prometheus가 타깃을 긁는 데 성공했는가를 0/1로 알려주는 **합성 지표**다.
|
||||
타깃이 응답하지 않으면 0이 된다.
|
||||
|
||||
**노드나 파드가 죽는 순간 가장 먼저 움직이는 신호**이며, 다른 모든 지표가
|
||||
사라지는 것과 달리 `up`은 **0이라는 값으로 남는다.** 그래서 "언제부터 죽었나"를
|
||||
사후에 알 수 있다.
|
||||
|
||||
### JGroups — 7800 차단 실험의 핵심
|
||||
|
||||
```promql
|
||||
vendor_jgroups_fd_sock2_get_num_suspected_members
|
||||
vendor_jgroups_merge3_get_views
|
||||
vendor_jgroups_tcp_get_different_cluster_messages
|
||||
```
|
||||
|
||||
| 지표 | 무엇을 말하는가 |
|
||||
|---|---|
|
||||
| `fd_sock2_..._suspected_members` | **FD_SOCK2가 의심하는 멤버 수.** 현재 두 파드 모두 `0`. 7800이 막히면 상대를 suspect 하기 시작한다 |
|
||||
| `merge3_get_views` | **MERGE3가 처리한 뷰 수.** split brain 후 다시 합칠 때 움직인다 |
|
||||
| `tcp_get_different_cluster_messages` | 다른 클러스터로부터 온 메시지 |
|
||||
|
||||
**7800 차단 실험의 가설** — `JGROUPS_PING` 테이블은 그대로 채워진 채
|
||||
`suspected_members`가 0에서 1로 오르고, 클러스터 뷰가 각각 1로 쪼개진다.
|
||||
|
||||
### 노드 지표
|
||||
|
||||
```promql
|
||||
node_memory_MemAvailable_bytes
|
||||
node_load1
|
||||
node_network_receive_bytes_total
|
||||
node_filesystem_avail_bytes
|
||||
```
|
||||
|
||||
**"머신이 죽었나 프로세스가 죽었나"** 를 가르는 데 쓴다. 파드는 사라졌는데
|
||||
node-exporter가 살아 있으면 프로세스 문제이고, 둘 다 사라지면 머신 문제다.
|
||||
|
||||
### Keycloak 애플리케이션 지표
|
||||
|
||||
```promql
|
||||
keycloak_session_expiration_task_seconds_count
|
||||
```
|
||||
|
||||
`keycloak_` 접두 지표는 아직 적다. 세션 관련 지표는 **실제 로그인이 발생해야**
|
||||
나타나므로, 세션 복제 실험 이후 다시 조사한다.
|
||||
|
||||
---
|
||||
|
||||
## 7. 접근
|
||||
|
||||
| | 주소 | 계정 |
|
||||
|---|---|---|
|
||||
| Grafana | `https://app2.hyeonworks.com` | `admin` / `lab-grafana-change-me` |
|
||||
| Prometheus | 클러스터 내부 `prometheus.observability.svc:9090` | — |
|
||||
|
||||
Prometheus UI를 직접 보려면 포트포워딩한다.
|
||||
|
||||
```bash
|
||||
kubectl -n observability port-forward svc/prometheus 9090:9090
|
||||
# http://localhost:9090/targets
|
||||
```
|
||||
|
||||
**Grafana 비밀번호가 매니페스트에 평문이다.** 로드맵 11번(비밀 관리)에서
|
||||
정리한다. 지금 드러내 두는 것은 의도이며, 감춰두면 잊어버린다.
|
||||
|
||||
---
|
||||
|
||||
## 8. 자원 실측
|
||||
|
||||
```
|
||||
grafana 65Mi
|
||||
prometheus 164Mi
|
||||
node-exporter 8Mi × 2
|
||||
────────────────────────
|
||||
합계 약 245Mi
|
||||
|
||||
kc-lab-1 2045Mi (41%)
|
||||
kc-lab-2 1131Mi (28%)
|
||||
호스트 여유 3957MB
|
||||
```
|
||||
|
||||
메모리 증설(8GB → 12GB) 전이었다면 kc-lab-1이 60%를 넘겼을 것이다.
|
||||
증설이 이 항목을 가능하게 했다.
|
||||
|
||||
---
|
||||
|
||||
## 9. 다음
|
||||
|
||||
관측이 서 있으므로 이제 고장을 주입하면 **무엇이 먼저 움직였는지**가 기록된다.
|
||||
|
||||
```
|
||||
0. 세션 복제 확인 ← 로그인 세션을 만들어 두 노드에 복제되는지
|
||||
1. TCP 7800 차단 ← suspected_members 와 JGROUPS_PING 대조
|
||||
2. DB 상실 ← postgres 파드 정지
|
||||
3. 노드 상실 ← kc-lab-2 (agent) 를 죽인다. kc-lab-1 이 아니다
|
||||
```
|
||||
|
||||
각 실험 전후로 같은 PromQL을 실행해 대조한다.
|
||||
|
||||
## 참고
|
||||
|
||||
| 문서 | 관계 |
|
||||
|---|---|
|
||||
| [`keycloak-multinode-cluster.md`](keycloak-multinode-cluster.md) | 관측 대상의 구성 |
|
||||
| [`session-lab-concepts.md`](session-lab-concepts.md) | Prometheus·RBAC·DaemonSet 등 개념 |
|
||||
| [`session-store-lab-roadmap.md`](session-store-lab-roadmap.md) | 로드맵 10번 |
|
||||
| [`two-hop-proxy-header-contract.md`](two-hop-proxy-header-contract.md) | `GF_SERVER_ROOT_URL`이 필요한 이유 |
|
||||
@@ -1,7 +1,15 @@
|
||||
# 열린 질문 커버리지 — 이 실험대로 답할 수 있는가
|
||||
|
||||
공개 기록(`hyeonworks.com/questions`)에 등록된 KeyCloak Patterns 열린 질문
|
||||
네 개를, 이 실험대가 실제로 검증할 수 있는지 대조한 결과.
|
||||
공개 기록에 등록된 KeyCloak Patterns 열린 질문 네 개를, 이 실험대가 실제로
|
||||
검증할 수 있는지 대조한 결과.
|
||||
|
||||
> **목록 경로는 [`/explore/questions`](https://hyeonworks.com/explore/questions)**
|
||||
> 다. `/questions` 는 404 이고 개별 문서만 `/questions/<slug>` 로 열린다.
|
||||
>
|
||||
> **2026-09-04 재확인** — Playwright 로 네 문서를 전문 재독하고
|
||||
> 「남은 미지수」·「다음 검증」·「제약」을 항목 단위로 대조한 결과
|
||||
> **계획에 빠진 항목 9개**를 찾아 보강했다. 항목별 실험 번호 대조표는
|
||||
> [`experiment-plan.md`](experiment-plan.md) B층 머리에 있다.
|
||||
|
||||
**결론 — 네 개 모두 이 실험대에서 재현 가능하다. 다만 로드맵에 빠진 항목이
|
||||
있고, 순서가 한 곳 뒤집혀 있다.**
|
||||
|
||||
@@ -2826,17 +2826,606 @@ SSH 공개키 두 줄이다. 공개키 자체는 비밀이 아니지만, **저
|
||||
|
||||
---
|
||||
|
||||
## 10층. 쿠버네티스 리소스 — 이 실험대에서 실제로 쓴 것들
|
||||
|
||||
5층이 k3s 자체라면 여기는 그 위에 올린 리소스들이다.
|
||||
|
||||
### 워크로드 세 종류 — 무엇을 언제 쓰는가
|
||||
|
||||
| | 보장하는 것 | 이 실험대에서 |
|
||||
|---|---|---|
|
||||
| **Deployment** | 파드 N개를 유지. 이름은 매번 바뀐다 | postgres, grafana, prometheus, echo |
|
||||
| **StatefulSet** | **안정된 이름**(`-0`, `-1`)과 순서 | **keycloak** |
|
||||
| **DaemonSet** | **노드마다 정확히 하나** | node-exporter, svclb |
|
||||
|
||||
**StatefulSet을 Keycloak에 쓴 이유** — Infinispan이 **파드 이름 + 랜덤 접미사**를
|
||||
클러스터 노드 식별자로 쓴다(`keycloak-0-49501`). Deployment면 이름이
|
||||
`keycloak-7d9f8b-x4k2p`처럼 매번 달라져서, 로그와 `JGROUPS_PING` 테이블을
|
||||
대조하기가 어려워진다.
|
||||
|
||||
**`podManagementPolicy`**
|
||||
|
||||
| 값 | 동작 |
|
||||
|---|---|
|
||||
| `OrderedReady` (기본) | `-0`이 Ready가 된 뒤에야 `-1`을 만든다 |
|
||||
| **`Parallel`** | **동시에 시작한다** |
|
||||
|
||||
이 실험대는 `Parallel`을 쓴다. 두 파드가 **동시에 클러스터 등록을 시도하는 것**이
|
||||
운영에서 실제로 일어나는 상황이기 때문이다.
|
||||
|
||||
**DaemonSet을 node-exporter에 쓴 이유** — replica 수를 지정하지 않는다.
|
||||
노드가 늘면 자동으로 늘고, 줄면 준다. **죽을 노드에도 반드시 있어야**
|
||||
꺼지기 직전의 마지막 샘플이 남는다.
|
||||
|
||||
```bash
|
||||
kubectl get deploy,sts,ds -A
|
||||
```
|
||||
|
||||
### 저장소 — PVC · PV · StorageClass
|
||||
|
||||
```
|
||||
PersistentVolumeClaim (PVC) "5Gi 짜리 읽기쓰기 볼륨을 주세요" ← 요청
|
||||
│ storageClassName: local-path
|
||||
▼
|
||||
StorageClass 어떻게 만들지 아는 프로비저너
|
||||
│
|
||||
▼
|
||||
PersistentVolume (PV) 실제로 만들어진 볼륨 ← 결과
|
||||
```
|
||||
|
||||
**PVC는 요청서, PV는 실물이다.** 파드는 PVC 이름만 알면 되고, 그 뒤가
|
||||
로컬 디스크인지 NFS인지 클라우드 블록 스토리지인지 몰라도 된다.
|
||||
|
||||
**`accessModes`**
|
||||
|
||||
| 값 | 의미 |
|
||||
|---|---|
|
||||
| **`ReadWriteOnce` (RWO)** | **한 노드에서만** 읽기/쓰기 |
|
||||
| `ReadOnlyMany` | 여러 노드에서 읽기만 |
|
||||
| `ReadWriteMany` | 여러 노드에서 읽기/쓰기 (NFS 등) |
|
||||
|
||||
**RWO가 `strategy: Recreate`를 강제한다.** 기본값 `RollingUpdate`는 새 파드를
|
||||
띄운 뒤 옛 파드를 내리는데, RWO 볼륨은 **두 파드가 동시에 마운트할 수 없어서**
|
||||
새 파드가 영원히 Pending에 머문다.
|
||||
|
||||
```yaml
|
||||
strategy:
|
||||
type: Recreate # 옛 파드를 먼저 내리고 새 파드를 띄운다
|
||||
```
|
||||
|
||||
**k3s의 `local-path` 프로비저너 — 볼륨이 노드에 못박힌다**
|
||||
|
||||
```json
|
||||
"nodeAffinity": {
|
||||
"required": { "nodeSelectorTerms": [{
|
||||
"matchExpressions": [{ "key": "kubernetes.io/hostname", "values": ["kc-lab-2"] }]
|
||||
}]}
|
||||
}
|
||||
경로: /var/lib/rancher/k3s/storage/pvc-<uuid>_<ns>_<name>
|
||||
```
|
||||
|
||||
**그 노드의 로컬 디스크에 디렉터리를 만드는 것이 전부**다. 따라서
|
||||
**PVC를 쓰는 파드는 그 노드를 벗어날 수 없다.**
|
||||
|
||||
| 결과 | |
|
||||
|---|---|
|
||||
| 노드가 죽으면 | **파드가 다른 노드로 재배치되지 못한다** |
|
||||
| 실험 관점 | **결함이 아니라 조건이다.** "DB가 있는 노드가 죽으면"이 의미를 갖는다 |
|
||||
|
||||
```bash
|
||||
kubectl get pvc -A
|
||||
kubectl get pv
|
||||
kubectl get pv <name> -o jsonpath='{.spec.nodeAffinity}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
### Secret — 감춰지지 않는다
|
||||
|
||||
```yaml
|
||||
kind: Secret
|
||||
type: Opaque
|
||||
stringData:
|
||||
POSTGRES_PASSWORD: lab-postgres-change-me
|
||||
```
|
||||
|
||||
`stringData`는 평문으로 쓰고 쿠버네티스가 base64로 인코딩해 저장한다.
|
||||
`data`는 직접 base64로 넣는다.
|
||||
|
||||
**base64는 암호화가 아니라 인코딩이다.**
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab get secret keycloak-lab-secrets -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
|
||||
```
|
||||
|
||||
한 줄로 읽힌다. etcd에도 그대로 들어 있다.
|
||||
|
||||
| 그래도 Secret을 쓰는 이유 | |
|
||||
|---|---|
|
||||
| RBAC로 접근을 나눌 수 있다 | ConfigMap과 별도로 권한 관리 |
|
||||
| 로그·`describe`에 값이 안 찍힌다 | 사고로 노출될 확률이 준다 |
|
||||
| 볼륨·env 주입 방식이 표준화된다 | |
|
||||
|
||||
**진짜 보호는 별도 계층이다** — SealedSecret, 외부 KMS, 또는 클라우드
|
||||
시크릿 매니저. 로드맵 11번의 주제다.
|
||||
|
||||
### RBAC — ServiceAccount · ClusterRole · Binding
|
||||
|
||||
Prometheus가 쿠버네티스 API에 물어서 타깃을 찾으려면 **읽기 권한**이 필요하다.
|
||||
|
||||
```
|
||||
ServiceAccount 파드가 쓰는 신원 (누구인가)
|
||||
│
|
||||
ClusterRoleBinding 신원과 권한을 잇는다
|
||||
│
|
||||
ClusterRole 무엇을 할 수 있는가 (리소스 × 동사)
|
||||
```
|
||||
|
||||
```yaml
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: [nodes, nodes/metrics, nodes/proxy, services, endpoints, pods]
|
||||
verbs: [get, list, watch]
|
||||
```
|
||||
|
||||
**`Role`과 `ClusterRole`의 차이** — `Role`은 한 네임스페이스 안에서만,
|
||||
`ClusterRole`은 클러스터 전체에서 유효하다. 노드는 네임스페이스에 속하지
|
||||
않으므로 **노드를 읽으려면 반드시 `ClusterRole`**이다.
|
||||
|
||||
**서브리소스가 따로 있다 — 실제로 걸린 함정**
|
||||
|
||||
`nodes`, `nodes/metrics`, `nodes/proxy`는 **서로 다른 권한**이다.
|
||||
|
||||
```
|
||||
/api/v1/nodes/<name>/proxy/metrics
|
||||
─────
|
||||
이 경로에는 nodes/proxy 가 필요
|
||||
```
|
||||
|
||||
`nodes/proxy`를 빠뜨렸을 때 kubelet 타깃만 **403 Forbidden**으로 실패하고
|
||||
나머지 잡은 전부 정상이었다. **부분 실패라 `rollout status`는 성공이라고
|
||||
말한다.** 타깃 목록을 직접 봐야 드러난다.
|
||||
|
||||
```bash
|
||||
kubectl auth can-i get nodes/proxy --as=system:serviceaccount:observability:prometheus
|
||||
kubectl describe clusterrole prometheus
|
||||
```
|
||||
|
||||
### 배치 제어 — nodeSelector · 라벨 · taint
|
||||
|
||||
```yaml
|
||||
nodeSelector:
|
||||
node-role.kubernetes.io/control-plane: "true"
|
||||
```
|
||||
|
||||
**호스트 이름 대신 역할 라벨을 쓴다.** `kubernetes.io/hostname: kc-lab-1`로
|
||||
못박으면 노드 이름이 바뀔 때 깨지고, **왜 거기 두는지가 드러나지 않는다.**
|
||||
|
||||
k3s는 server 노드에 `node-role.kubernetes.io/control-plane=true`를 붙인다.
|
||||
|
||||
```bash
|
||||
kubectl get nodes --show-labels
|
||||
kubectl get nodes -l node-role.kubernetes.io/control-plane=true
|
||||
```
|
||||
|
||||
**taint와 toleration**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **taint** | 노드에 붙는 "여기 오지 마" 표시 |
|
||||
| **toleration** | 파드가 갖는 "그래도 갈 수 있음" 면제권 |
|
||||
|
||||
```yaml
|
||||
tolerations:
|
||||
- operator: Exists # 어떤 taint 든 무시한다
|
||||
```
|
||||
|
||||
node-exporter에 이걸 주는 이유는 **관측이 빠지는 노드가 있으면 안 되기**
|
||||
때문이다. taint가 걸린 노드에서도 떠야 한다.
|
||||
|
||||
**배치를 정하는 세 수단의 차이**
|
||||
|
||||
| 수단 | 성격 |
|
||||
|---|---|
|
||||
| `nodeSelector` | **반드시** 그 라벨의 노드에 |
|
||||
| `topologySpreadConstraints` | **골고루** 퍼뜨린다 |
|
||||
| taint / toleration | 노드가 **거부**하고 파드가 **면제**받는다 |
|
||||
|
||||
### k3s server와 agent — 죽였을 때가 다르다
|
||||
|
||||
```bash
|
||||
kubectl get nodes -o custom-columns=\
|
||||
'NODE:.metadata.name,CP:.metadata.labels.node-role\.kubernetes\.io/control-plane'
|
||||
```
|
||||
|
||||
| | kc-lab-1 (**server**) | kc-lab-2 (**agent**) |
|
||||
|---|---|---|
|
||||
| 실행 | API 서버 · 스케줄러 · etcd(SQLite) | kubelet · containerd |
|
||||
| 이 실험대에서 | keycloak-1 · traefik · **coredns** · metrics-server · local-path-provisioner | keycloak-0 · postgres |
|
||||
| 죽이면 | **`kubectl`이 안 된다. DNS·인그레스도 사라진다** | 클러스터 제어는 살아 있다 |
|
||||
|
||||
**노드 상실 실험은 agent를 죽이는 것이다.** server를 죽이는 것은 노드 상실이
|
||||
아니라 **컨트롤 플레인 상실**이며 성격이 완전히 다르다.
|
||||
|
||||
이 사실을 모르고 "keycloak 하나만 있는 노드를 죽이자"고 계획했다가
|
||||
실제 배치를 조회한 뒤 정정했다.
|
||||
|
||||
---
|
||||
|
||||
## 11층. Keycloak 클러스터링 내부 — Infinispan과 JGroups
|
||||
|
||||
### 두 층으로 되어 있다
|
||||
|
||||
```
|
||||
Infinispan 분산 캐시. "세션을 어디에 두고 어떻게 복제할까"
|
||||
│
|
||||
JGroups 그룹 통신. "누가 멤버이고 어떻게 메시지를 주고받을까"
|
||||
│
|
||||
TCP 7800 실제 소켓
|
||||
```
|
||||
|
||||
Keycloak은 Infinispan을 쓰고, Infinispan은 JGroups 위에서 돈다.
|
||||
로그의 `org.infinispan.CLUSTER`와 `vendor_jgroups_*` 지표가 각각 이 두 층이다.
|
||||
|
||||
### 디스커버리와 트랜스포트는 다른 경로다
|
||||
|
||||
**이것이 이 실험대를 2노드로 만든 이유다.**
|
||||
|
||||
| 단계 | 경로 | 끊기면 |
|
||||
|---|---|---|
|
||||
| **디스커버리** — 서로를 찾는다 | PostgreSQL `JGROUPS_PING` 테이블 | 상대의 존재를 모른다 |
|
||||
| **트랜스포트** — 실제로 대화한다 | **TCP 7800** | **DB엔 등록되는데 클러스터가 안 붙는다** |
|
||||
|
||||
`JGROUPS_PING` 한 테이블에 두 메커니즘이 다 보인다.
|
||||
|
||||
```
|
||||
name | cluster_name | ip | coord
|
||||
------------------+--------------+-----------------+-------
|
||||
keycloak-0-49501 | ISPN | 10.42.1.18:7800 | f
|
||||
keycloak-1-26938 | ISPN | 10.42.0.16:7800 | t
|
||||
───────────────────────────── ──── ─
|
||||
디스커버리 결과 트랜스포트 경로 코디네이터
|
||||
```
|
||||
|
||||
전체 스키마는 `address / name / cluster_name / ip / coord / last_update /
|
||||
coordinated_by`이고 기본키는 `address`다.
|
||||
|
||||
> 오래된 자료에는 `own_addr`, `ping_data` 같은 컬럼명이 나오지만 Keycloak 26의
|
||||
> 실제 스키마는 위와 같다. 쿼리 전에 `\d jgroups_ping`으로 확인한다.
|
||||
|
||||
**`jdbc-ping`을 쓰는 이유** — 예전에는 UDP 멀티캐스트로 서로를 찾았다.
|
||||
쿠버네티스나 클라우드에서는 멀티캐스트가 막혀 있는 경우가 많아,
|
||||
**이미 있는 데이터베이스를 게시판처럼 쓰는** 방식으로 바뀌었다.
|
||||
Keycloak 26의 기본값이다.
|
||||
|
||||
### 코디네이터
|
||||
|
||||
`coord = t` 인 노드가 **코디네이터**다. 뷰 변경을 확정하고 리밸런싱을
|
||||
주도한다. 특별한 권한이 아니라 **역할**이며, 그 노드가 사라지면 남은 멤버가
|
||||
인계받는다.
|
||||
|
||||
실험대를 전원 종료했다 켰을 때 코디네이터가 `keycloak-1` → `keycloak-0`으로
|
||||
바뀌는 것을 관찰했다. **먼저 뜬 쪽이 맡는다.**
|
||||
|
||||
### 클러스터 뷰
|
||||
|
||||
```
|
||||
ISPN000094: Received new cluster view for channel ISPN:
|
||||
[keycloak-1-26938(v=16.0.12)|1] (2) [keycloak-1-26938, keycloak-0-49501]
|
||||
─────────────────────────── ─ ─ ────────────────────────────────────
|
||||
뷰를 만든 코디네이터 뷰 ID 멤버 수 멤버 목록
|
||||
```
|
||||
|
||||
**뷰(view)는 "지금 이 순간의 멤버 명단"** 이다. 멤버가 들어오거나 나가면
|
||||
새 뷰가 발행되고 뷰 ID가 올라간다.
|
||||
|
||||
| 로그 코드 | 의미 |
|
||||
|---|---|
|
||||
| `ISPN000094` | 새 클러스터 뷰를 받았다 |
|
||||
| `ISPN000079` | 자기 주소와 물리 주소(7800) |
|
||||
| `ISPN100000` | 노드가 합류했다 |
|
||||
|
||||
```bash
|
||||
kubectl -n keycloak-lab logs keycloak-0 | grep -E 'ISPN000094|ISPN000079|ISPN100000'
|
||||
```
|
||||
|
||||
### 주요 JGroups 프로토콜 — 지표 이름에 그대로 나온다
|
||||
|
||||
| 프로토콜 | 하는 일 | 관련 지표 |
|
||||
|---|---|---|
|
||||
| **GMS** (Group Membership Service) | 멤버십 관리, 뷰 발행 | `vendor_jgroups_gms_*` |
|
||||
| **FD_SOCK2** (Failure Detection) | **TCP 소켓으로 상대 생존 감시** | `..._get_num_suspected_members` |
|
||||
| **MERGE3** | **split brain 후 다시 합치기** | `..._merge3_get_views` |
|
||||
| **NAKACK2** | 신뢰성 있는 메시지 전달, 재전송 | `..._nakack2_*` |
|
||||
| **TCP** | 트랜스포트 | `..._tcp_*` |
|
||||
|
||||
**7800을 막으면 FD_SOCK2가 먼저 반응한다.** 소켓 연결이 끊기면 상대를
|
||||
suspect 하고, GMS가 그 멤버를 뷰에서 제외한다. 각자 자기만 있는 뷰가 되면
|
||||
**split brain**이고, 통신이 복구되면 MERGE3가 합친다.
|
||||
|
||||
### 세션은 어디에 있는가 — 두 곳이되 역할이 다르다
|
||||
|
||||
Keycloak 26의 기본값 `persistent-user-sessions`에서는
|
||||
|
||||
| 저장소 | 역할 | 노드 간 공유 |
|
||||
|---|---|---|
|
||||
| **PostgreSQL** | **진실의 원천.** 재시작에도 살아남는다 | **여기서만 일어난다** |
|
||||
| **Infinispan `sessions`** | **자기 노드가 로그인시킨 세션만** 담는 룩어사이드 캐시 | **일어나지 않는다** |
|
||||
|
||||
> **처음에 이 표에 "Infinispan = 캐시 + 노드 간 실시간 전파"라고 썼는데
|
||||
> 틀렸다.** 실험 0에서 측정해보니 세션 엔트리는 노드 사이를 건너가지 않는다.
|
||||
> 두 노드가 같은 답을 하는 이유는 복제가 아니라 같은 DB를 보기 때문이고,
|
||||
> 반대편 노드가 실제로 날리는 `SELECT ... FROM OFFLINE_USER_SESSION` 을
|
||||
> PostgreSQL 로그에서 직접 잡았다.
|
||||
> → [`docs/experiment-00-session-replication.md`](experiment-00-session-replication.md)
|
||||
|
||||
`--features-disabled=persistent-user-sessions`로 끄면 Infinispan만 남는
|
||||
**volatile** 모드가 되고, 그때는 캐시가 곧 진실의 원천이므로 **복제가
|
||||
반드시 일어나야 한다.** 이 둘의 차이가 로드맵 2번의 주제다.
|
||||
|
||||
### 세션 쓰기 트랜잭션의 세 가지 설계 결정
|
||||
|
||||
PostgreSQL 문장 로깅으로 잡은 갱신 트랜잭션 하나에 다 들어 있다.
|
||||
|
||||
| 보이는 것 | 뜻 |
|
||||
|---|---|
|
||||
| `update ... where ... and VERSION=$5` | **낙관적 락.** 읽을 때의 버전과 같을 때만 쓴다 |
|
||||
| `for no key update ... skip locked` | 잠긴 행을 **기다리지 않고 건너뛴다.** 대기 대신 재시도 |
|
||||
| **`SET LOCAL synchronous_commit TO OFF`** | **WAL 플러시를 기다리지 않고 커밋한다** |
|
||||
|
||||
마지막 것이 특히 중요하다 — **DB가 강제 종료되면 직전 수백 밀리초의 세션
|
||||
갱신이 사라질 수 있다.** 버그가 아니라 의도된 트레이드오프다.
|
||||
`LAST_SESSION_REFRESH` 갱신은 매우 잦고, 잃어도 사용자가 다시 갱신하면 된다.
|
||||
|
||||
---
|
||||
|
||||
## 12층. 관측성 — Prometheus의 구조
|
||||
|
||||
### 세 부분으로 되어 있다
|
||||
|
||||
```
|
||||
수집(scrape) ──▶ 저장(TSDB) ──▶ 질의(PromQL)
|
||||
15초마다 로컬 디스크 Grafana 또는 API
|
||||
HTTP GET /metrics 시계열
|
||||
```
|
||||
|
||||
**Prometheus는 pull 방식이다.** 대상이 보내주는 것이 아니라 Prometheus가
|
||||
주기적으로 `/metrics`를 긁어간다.
|
||||
|
||||
| 결과 | |
|
||||
|---|---|
|
||||
| 대상이 죽으면 | 긁기가 실패하고 **`up`이 0이 된다** — 죽은 사실 자체가 데이터가 된다 |
|
||||
| 방화벽 방향 | Prometheus → 대상. 대상이 Prometheus 주소를 알 필요가 없다 |
|
||||
| 짧은 작업 | 긁히기 전에 끝나면 잡히지 않는다 (Pushgateway가 필요한 경우) |
|
||||
|
||||
### exporter 패턴
|
||||
|
||||
애플리케이션이 Prometheus 형식을 모를 때, **번역기**를 옆에 둔다.
|
||||
|
||||
| exporter | 무엇을 노출하는가 |
|
||||
|---|---|
|
||||
| **node-exporter** | 머신 — CPU, 메모리, 디스크, 네트워크 |
|
||||
| kube-state-metrics | 쿠버네티스 오브젝트 상태 |
|
||||
| postgres-exporter | PostgreSQL 내부 통계 |
|
||||
|
||||
**Keycloak과 Traefik은 exporter가 필요 없다.** 자체적으로 Prometheus 형식
|
||||
엔드포인트를 제공한다(`KC_METRICS_ENABLED=true`).
|
||||
|
||||
### 서비스 디스커버리 — 타깃을 적어두지 않는다
|
||||
|
||||
```yaml
|
||||
kubernetes_sd_configs:
|
||||
- role: endpoints
|
||||
namespaces: { names: [keycloak-lab] }
|
||||
```
|
||||
|
||||
**파드 IP는 재시작마다 바뀐다.** 실험대를 전원 종료했다 켜니 모든 파드가
|
||||
새 주소를 받았다(`10.42.1.22` → `10.42.1.25`). 정적 목록은 그때마다 깨진다.
|
||||
|
||||
`role`에 따라 무엇을 찾을지가 달라진다.
|
||||
|
||||
| role | 찾는 것 |
|
||||
|---|---|
|
||||
| `endpoints` | 서비스 뒤의 실제 파드들 ← 애플리케이션 지표 |
|
||||
| `node` | 노드 |
|
||||
| `pod` | 파드 직접 |
|
||||
| `service` | 서비스 |
|
||||
|
||||
### relabel — 걸러내고 이름을 붙인다
|
||||
|
||||
디스커버리는 **전부 다** 가져온다. 그중 필요한 것만 남기는 것이 relabel이다.
|
||||
|
||||
```yaml
|
||||
relabel_configs:
|
||||
- source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
|
||||
action: keep
|
||||
regex: keycloak-headless;management
|
||||
- source_labels: [__meta_kubernetes_pod_name]
|
||||
target_label: pod
|
||||
```
|
||||
|
||||
| `action` | 하는 일 |
|
||||
|---|---|
|
||||
| `keep` | regex에 맞는 것만 남긴다 |
|
||||
| `drop` | 맞는 것을 버린다 |
|
||||
| `replace` (기본) | 라벨 값을 만든다 |
|
||||
| `labelmap` | 메타 라벨을 일반 라벨로 복사 |
|
||||
|
||||
**`__`로 시작하는 라벨은 내부용**이며 저장되지 않는다. `__meta_*`는
|
||||
디스커버리가 붙여준 정보이고, 필요하면 `target_label`로 옮겨야 남는다.
|
||||
|
||||
**`pod`과 `node` 라벨을 붙이는 것이 실험에서 결정적이다.** 없으면
|
||||
"어느 파드가, 어느 노드에서"에 답할 수 없다.
|
||||
|
||||
### 메트릭 타입
|
||||
|
||||
| 타입 | 성질 | 예 |
|
||||
|---|---|---|
|
||||
| **counter** | **누적. 줄지 않는다** (재시작 시 0으로) | `..._requests_total` |
|
||||
| **gauge** | 오르내린다 | `node_memory_MemAvailable_bytes` |
|
||||
| **histogram** | 구간별 분포 + 합계 + 개수 | `..._seconds_bucket/_sum/_count` |
|
||||
| summary | 분위수를 클라이언트가 계산 | |
|
||||
|
||||
**counter는 그대로 보면 의미가 없다.** 변화율을 봐야 한다.
|
||||
|
||||
```promql
|
||||
rate(http_requests_total[5m])
|
||||
```
|
||||
|
||||
**histogram은 세 지표가 한 벌**이다. `_bucket`으로 분위수를 계산한다.
|
||||
|
||||
```promql
|
||||
histogram_quantile(0.95, rate(keycloak_session_expiration_task_seconds_bucket[5m]))
|
||||
```
|
||||
|
||||
### `up` — 가장 중요한 합성 지표
|
||||
|
||||
```promql
|
||||
up
|
||||
up{job="keycloak"}
|
||||
```
|
||||
|
||||
Prometheus가 **직접 만드는** 지표다. 긁기에 성공하면 1, 실패하면 0.
|
||||
|
||||
**장애 실험에서 이것이 핵심인 이유** — 다른 지표는 대상이 죽으면 **사라진다.**
|
||||
사라진 데이터로는 "언제부터 죽었나"를 알 수 없다. `up`은 **0이라는 값으로
|
||||
남기 때문에** 사후에 시각을 특정할 수 있다.
|
||||
|
||||
```promql
|
||||
up == 0 # 지금 죽은 타깃
|
||||
changes(up[1h]) # 1시간 동안 몇 번 오르내렸나
|
||||
min_over_time(up[10m]) # 10분 중 한 번이라도 죽었나
|
||||
```
|
||||
|
||||
### TSDB와 보존 기간
|
||||
|
||||
```yaml
|
||||
--storage.tsdb.path=/prometheus
|
||||
--storage.tsdb.retention.time=7d
|
||||
```
|
||||
|
||||
로컬 디스크에 시계열로 저장한다. **보존 기간이 지나면 삭제**되므로 볼륨이
|
||||
무한히 커지지 않는다.
|
||||
|
||||
`emptyDir`에 두면 파드 재시작 시 **실험 기록이 통째로 사라진다.**
|
||||
사후 추적이 목적이면 PVC여야 한다.
|
||||
|
||||
### 관측 시스템의 장애 도메인
|
||||
|
||||
**관측 시스템은 관측 대상과 같이 죽으면 안 된다.** 죽는 순간을 기록해야
|
||||
하는데 같이 죽으면 기록이 없다.
|
||||
|
||||
노드가 둘뿐인 실험대에서는 완전히 피할 수 없으므로 **규칙으로 정한다.**
|
||||
|
||||
```
|
||||
kc-lab-1 (server) 관측 스택을 둔다. 죽이지 않는다
|
||||
kc-lab-2 (agent) 장애 주입 대상
|
||||
```
|
||||
|
||||
`nodeSelector`로 못박아 실험이 재현 가능하게 만든다.
|
||||
|
||||
---
|
||||
|
||||
## 13층. 가상화 운영 — 실행 중 바꾸는 것들
|
||||
|
||||
### VM 메모리 재배분 — 게스트를 다시 만들지 않는다
|
||||
|
||||
```bash
|
||||
virsh setmaxmem kc-lab-1 5120M --config
|
||||
virsh setmem kc-lab-1 5120M --config
|
||||
```
|
||||
|
||||
| 명령 | 바꾸는 것 |
|
||||
|---|---|
|
||||
| `setmaxmem` | **상한**. 부팅 시 게스트가 보는 총량 |
|
||||
| `setmem` | **현재 할당**. 상한 이하여야 한다 |
|
||||
|
||||
**순서가 중요하다.** 현재값을 상한보다 크게 줄 수 없으므로 `setmaxmem`이
|
||||
먼저다.
|
||||
|
||||
| 플래그 | 적용 범위 |
|
||||
|---|---|
|
||||
| `--config` | 영구 정의. **다음 부팅부터** |
|
||||
| `--live` | 실행 중인 도메인에 즉시 |
|
||||
| 둘 다 | 지금과 앞으로 |
|
||||
|
||||
`setmaxmem --live`는 대개 거부된다 — 게스트가 부팅 시 메모리 맵을 정하기
|
||||
때문이다. **상한을 바꾸려면 게스트를 껐다 켜야 한다.**
|
||||
|
||||
```bash
|
||||
virsh dominfo kc-lab-1 | grep -i memory
|
||||
ssh kc-lab-1 free -m # 게스트가 실제로 인식한 값
|
||||
```
|
||||
|
||||
호스트에서 8GB→12GB로 물리 증설한 뒤 이 방법으로 재배분했다.
|
||||
**게스트 재생성이나 디스크 조작은 전혀 필요 없었다.**
|
||||
|
||||
### 안전한 종료 순서
|
||||
|
||||
전원을 내리기 전에 **위에서부터** 정리한다.
|
||||
|
||||
```bash
|
||||
# 1. 애플리케이션 — 클러스터에서 정상 탈퇴
|
||||
kubectl -n keycloak-lab scale statefulset/keycloak --replicas=0
|
||||
kubectl -n keycloak-lab wait --for=delete pod -l app=keycloak --timeout=120s
|
||||
|
||||
# 2. 데이터베이스 — 마지막에, 충분한 시간을 주고
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=0
|
||||
kubectl -n keycloak-lab wait --for=delete pod -l app=postgres --timeout=120s
|
||||
|
||||
# 3. 게스트 — ACPI 정상 종료
|
||||
virsh shutdown kc-lab-1 && virsh shutdown kc-lab-2
|
||||
|
||||
# 4. 호스트
|
||||
sudo systemctl poweroff
|
||||
```
|
||||
|
||||
**왜 순서가 중요한가** — `virsh shutdown`은 게스트 systemd가 k3s를 멈추고,
|
||||
k3s가 컨테이너에 SIGTERM을 보낸다. 유예 시간이 짧으면 **PostgreSQL이
|
||||
강제 종료되어 다음 기동에 crash recovery가 돈다.** 미리 내려두면 그 위험이
|
||||
없다.
|
||||
|
||||
**clean shutdown 확인**
|
||||
|
||||
```bash
|
||||
ssh kc-lab-2 'sudo ls /var/lib/rancher/k3s/storage/*postgres-data*/pgdata/postmaster.pid'
|
||||
```
|
||||
|
||||
**`postmaster.pid`가 남아 있지 않아야 정상**이다. 남아 있으면 비정상 종료였고
|
||||
다음 기동에 복구 절차가 실행된다.
|
||||
|
||||
### 복구 순서 — 종료의 역순
|
||||
|
||||
```bash
|
||||
virsh start kc-lab-1 && virsh start kc-lab-2
|
||||
kubectl get nodes # Ready 2개 대기
|
||||
kubectl -n keycloak-lab scale deployment/postgres --replicas=1
|
||||
kubectl -n keycloak-lab rollout status deployment/postgres
|
||||
kubectl -n keycloak-lab scale statefulset/keycloak --replicas=2
|
||||
```
|
||||
|
||||
**PostgreSQL이 먼저다.** Keycloak이 DB 없이 뜨면 기동에 실패한다.
|
||||
|
||||
**스케일을 0으로 내려두면 자동으로 복구되지 않는다.** 명시적으로 올려야 한다.
|
||||
|
||||
---
|
||||
|
||||
## 아직 기록하지 않은 개념
|
||||
|
||||
실험 설계 단계에서 아래 항목을 이 문서에 추가한다.
|
||||
실험을 진행하면서 이 문서에 추가한다.
|
||||
|
||||
- Infinispan, `DIST_SYNC`, `numOwners`, 캐시별 설정
|
||||
- JGroups, `JDBC_PING`, 디스커버리와 트랜스포트의 분리, TCP 7800
|
||||
- `persistent-user-sessions` / `volatile-user-sessions`
|
||||
- 원격 Infinispan(Hot Rod)과 multi-site
|
||||
- refresh token rotation, revoke, max reuse, 동시 갱신 경쟁
|
||||
- `persistent-user-sessions` / `volatile-user-sessions` 의 실제 차이 (로드맵 2번)
|
||||
- refresh token rotation·revoke·max reuse 와 동시 갱신 경쟁 (로드맵 5번)
|
||||
- SSO 세션 vs 애플리케이션 세션, `KEYCLOAK_IDENTITY`, `AUTH_SESSION_ID`
|
||||
- 백채널 로그아웃과 `sid` 역인덱스
|
||||
- 쿠키 `Secure` / `SameSite` / `HttpOnly`
|
||||
- Redis 영속화(RDB/AOF)와 세션 복구
|
||||
- `tc netem`, OOM killer와 `oom_score`, fsync와 페이지 캐시
|
||||
- Spring Session / `OAuth2AuthorizedClientService` 의 저장 구조
|
||||
- `tc netem` 지연 주입
|
||||
- OOM killer 와 `oom_score`
|
||||
- fsync 와 페이지 캐시, EBS IOPS
|
||||
|
||||
### 이번에 채운 것 (2026-09-04)
|
||||
|
||||
10~13층으로 기록 완료 — StatefulSet·DaemonSet, PVC/PV/StorageClass,
|
||||
Secret, RBAC 와 서브리소스, nodeSelector·taint, k3s server/agent 차이,
|
||||
Infinispan·JGroups(디스커버리 vs 트랜스포트, GMS/FD_SOCK2/MERGE3),
|
||||
Prometheus(pull·SD·relabel·메트릭 타입·`up`·TSDB), VM 메모리 재배분,
|
||||
안전한 종료·복구 순서.
|
||||
|
||||
@@ -0,0 +1,400 @@
|
||||
# 이 실험을 이해하기 위한 선수 지식
|
||||
|
||||
실험 결과를 먼저 들이밀었더니 맥락이 사라졌다. 이 문서는 **왜 이런 걸
|
||||
측정하고 있는지**를 바닥부터 세운다.
|
||||
|
||||
읽는 순서가 곧 의존 관계다. 아는 절은 건너뛰어도 되지만, 3장까지는
|
||||
"세션"이라는 말의 뜻이 계속 바뀌므로 훑고 가는 편이 낫다.
|
||||
|
||||
---
|
||||
|
||||
## 0. 출발점 — 당신이 원래 물은 것
|
||||
|
||||
> Keycloak이 여러 개일 경우, 세션 저장소를 Redis나 별도 저장소로 쓸 경우,
|
||||
> Redis와 DB에 분리해서 세션과 토큰을 관리할 때 어떻게 달라지는지.
|
||||
> SSO를 추가하면 어떻게 달라지는지. Redis 또는 DB가 죽으면 어떻게 복구하는지.
|
||||
|
||||
이 질문에 답하려면 **"세션이 어디에 있는가"** 를 정확히 알아야 한다.
|
||||
지금 하고 있는 실험은 전부 그 한 문장을 쪼갠 것이다.
|
||||
|
||||
---
|
||||
|
||||
## 1. HTTP는 기억이 없다
|
||||
|
||||
모든 것의 출발점.
|
||||
|
||||
```
|
||||
요청 1: GET /login → 서버
|
||||
요청 2: GET /mypage → 서버 ← 서버는 요청 1을 기억하지 못한다
|
||||
```
|
||||
|
||||
HTTP 요청은 **하나하나가 완전히 독립적**이다. 서버 입장에서 두 번째 요청은
|
||||
생판 처음 보는 사람이 보낸 것과 구별되지 않는다.
|
||||
|
||||
그래서 "로그인했다"는 사실을 **어딘가 저장**해야 한다.
|
||||
|
||||
```
|
||||
브라우저 서버
|
||||
┌──────────────┐ ┌────────────────────────┐
|
||||
│ 쿠키 │ │ 세션 저장소 │
|
||||
│ SESSIONID= │ ──── 매 요청 ────▶ │ abc123 → { │
|
||||
│ abc123 │ 이 값만 보냄 │ user: "홍길동", │
|
||||
└──────────────┘ │ 로그인시각: ... │
|
||||
│ } │
|
||||
작은 표만 들고 다닌다 └────────────────────────┘
|
||||
실제 내용은 여기 있다
|
||||
```
|
||||
|
||||
| 용어 | 뜻 |
|
||||
|---|---|
|
||||
| **쿠키** | 브라우저가 들고 다니는 **작은 표(번호표)**. 보통 세션 ID만 들어 있다 |
|
||||
| **세션** | 서버가 그 번호에 대해 기억하는 **실제 내용** |
|
||||
|
||||
**서버가 1대면 여기서 이야기가 끝난다.** 문제는 2대부터다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 서버가 2대가 되는 순간 — 이 실험의 진짜 출발점
|
||||
|
||||
```
|
||||
로그인 요청 ──▶ 서버 A A의 메모리에 "abc123 = 홍길동" 기록
|
||||
다음 요청 ──▶ 서버 B B: "abc123? 그런 거 모르는데" → 로그아웃 화면
|
||||
```
|
||||
|
||||
**이게 전부다.** 분산 세션이라는 주제 전체가 이 한 장면에서 나온다.
|
||||
|
||||
푸는 방법은 셋뿐이다.
|
||||
|
||||
| 방법 | 어떻게 | 대가 |
|
||||
|---|---|---|
|
||||
| **1. 고정 배정** (sticky session) | 같은 사람은 항상 같은 서버로 보낸다 | **그 서버가 죽으면 그 사람 세션은 사라진다.** 부하도 안 고르게 퍼진다 |
|
||||
| **2. 복제** | 서버끼리 메모리 내용을 서로 보낸다 | 서버가 N대면 트래픽이 N² 로 는다. 어긋남(불일치)이 생긴다 |
|
||||
| **3. 공유 저장소** | 세션을 바깥(DB·Redis)에 두고 모두가 본다 | **그게 죽으면 전체가 멈춘다.** 매 요청마다 네트워크 왕복 |
|
||||
|
||||
> **당신이 원래 물은 "Redis나 별도 저장소를 쓰면"이 바로 3번**이다.
|
||||
> 그리고 "Redis나 DB가 죽으면 어떻게 복구하나"는 3번의 대가를 묻는 것이다.
|
||||
|
||||
**Keycloak도 예외가 아니다.** Keycloak을 2대 띄우면 정확히 이 문제가 생긴다.
|
||||
Keycloak이 이걸 어떻게 풀었는지가 실험 0의 주제다.
|
||||
|
||||
---
|
||||
|
||||
## 3. Keycloak은 무엇이고, 왜 세션을 갖는가
|
||||
|
||||
### 3-1. 하는 일
|
||||
|
||||
Keycloak은 **로그인을 대신 해주는 서버**다.
|
||||
|
||||
```
|
||||
[사용자] [내 앱] [Keycloak]
|
||||
│ │ │
|
||||
│─ 접속 ──────▶│ │
|
||||
│◀─ "Keycloak 가서 로그인하고 와" ────│
|
||||
│──────────────────── 로그인 ───────▶│
|
||||
│◀─────────────── 토큰 발급 ─────────│
|
||||
│─ 토큰 들고 ──▶│ │
|
||||
│ │─ 이 토큰 유효해? ──▶│
|
||||
```
|
||||
|
||||
내 앱은 비밀번호를 저장하지도, 검증하지도 않는다. 그 일을 Keycloak이 한다.
|
||||
|
||||
### 3-2. 그래서 **세션이 두 겹**이 된다
|
||||
|
||||
여기가 헷갈리는 지점이다. "세션"이라는 말이 두 가지를 가리킨다.
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────┐
|
||||
│ Keycloak 의 SSO 세션 │
|
||||
│ "이 브라우저는 홍길동으로 로그인되어 있다" │
|
||||
│ 쿠키 이름: KEYCLOAK_IDENTITY │
|
||||
└─────────────────────────────────────────────────────┘
|
||||
│ │
|
||||
▼ ▼
|
||||
┌──────────────────┐ ┌──────────────────┐
|
||||
│ 앱1 의 세션 │ │ 앱2 의 세션 │
|
||||
│ (또는 토큰) │ │ (또는 토큰) │
|
||||
└──────────────────┘ └──────────────────┘
|
||||
```
|
||||
|
||||
| | 누가 갖는가 | 사라지면 |
|
||||
|---|---|---|
|
||||
| **SSO 세션** | **Keycloak** | 모든 앱에서 다시 로그인해야 한다 |
|
||||
| 앱 세션 | 각 애플리케이션 | 그 앱만 다시 들어가면 된다 |
|
||||
|
||||
**SSO가 되는 원리가 이것이다.** 앱1에서 로그인하면 Keycloak에 SSO 세션이
|
||||
생긴다. 앱2로 가면 Keycloak이 "이 브라우저 이미 로그인했네" 하고 **로그인
|
||||
화면 없이** 바로 토큰을 준다.
|
||||
|
||||
> **그래서 Keycloak의 세션이 사라지면 SSO 전체가 깨진다.**
|
||||
> 당신 질문의 "SSO를 추가하면 어떻게 달라지는지"가 여기 걸린다.
|
||||
> 앱이 하나일 때는 그 앱만 재로그인이지만, SSO에서는 **전 앱이 동시에** 터진다.
|
||||
|
||||
---
|
||||
|
||||
## 4. 토큰이 있는데 왜 세션이 필요한가
|
||||
|
||||
가장 흔한 오해다. "JWT는 stateless라서 서버가 기억할 게 없다"는 말은
|
||||
**반만 맞다.**
|
||||
|
||||
### 4-1. 토큰이 두 종류다
|
||||
|
||||
```
|
||||
로그인 성공
|
||||
│
|
||||
├──▶ access token 수명 짧음 (이 실험대: 60초)
|
||||
│ JWT. 서명이 붙어 있어 서버가 아무것도 기억 안 해도 검증된다
|
||||
│ → 진짜 stateless
|
||||
│
|
||||
└──▶ refresh token 수명 김 (이 실험대: 1800초 = 30분)
|
||||
access token 이 만료되면 이걸로 새로 받는다
|
||||
→ 서버가 세션을 기억하고 있어야 한다
|
||||
```
|
||||
|
||||
| | access token | refresh token |
|
||||
|---|---|---|
|
||||
| 검증 방식 | **서명만 보면 됨** | **서버 세션 조회 필요** |
|
||||
| 취소 | **불가능** (만료를 기다려야) | 가능 |
|
||||
| 수명 | 짧게 (분 단위) | 길게 (시간~일) |
|
||||
|
||||
### 4-2. 그래서 이렇게 된다
|
||||
|
||||
```
|
||||
0초 로그인 세션 생성
|
||||
0초 access token 발급 이후 60초간은 서버에 안 물어봐도 됨
|
||||
60초 access token 만료
|
||||
60초 refresh 요청 ─────▶ 서버: "이 세션 살아 있나?" ◀── 여기서 세션 필요
|
||||
60초 새 access token
|
||||
120초 또 만료 → 또 refresh → 또 세션 조회
|
||||
```
|
||||
|
||||
**60초마다 세션 저장소를 친다.** access token 수명이 짧을수록 세션 저장소
|
||||
부하가 커진다 — 보안과 성능의 맞바꿈이 여기서 일어난다.
|
||||
|
||||
### 4-3. 로그아웃도 세션이 있어야 한다
|
||||
|
||||
**로그아웃 = 세션 삭제**다. 세션이 없으면 로그아웃이라는 개념 자체가 없다.
|
||||
이미 발급된 access token은 서명이 유효하므로 만료 전까지 계속 통과한다.
|
||||
|
||||
> 그래서 access token 수명을 60초로 짧게 잡는다. 로그아웃해도 최대 60초는
|
||||
> 살아 있다는 뜻이고, 그 이상은 refresh 가 막히므로 끝난다.
|
||||
|
||||
---
|
||||
|
||||
## 5. Keycloak은 세션을 어디에 두는가 — **버전에 따라 답이 다르다**
|
||||
|
||||
이 실험 전체가 여기에 걸려 있다.
|
||||
|
||||
### 5-1. 두 개의 후보
|
||||
|
||||
| | 무엇 | 성질 |
|
||||
|---|---|---|
|
||||
| **Infinispan** | Keycloak **안에 내장된** 분산 캐시. Java 라이브러리 | 메모리. 빠름. 프로세스가 죽으면 사라짐 |
|
||||
| **데이터베이스** | PostgreSQL 등 바깥의 DB | 디스크. 느림. 재시작해도 남음 |
|
||||
|
||||
**Infinispan은 별도로 설치하는 물건이 아니다.** Keycloak 프로세스 안에서 도는
|
||||
라이브러리다. Redis처럼 따로 띄우는 게 아니다 — 이걸 헷갈리면 전체가 안 맞는다.
|
||||
|
||||
### 5-2. 버전별로 이렇게 바뀌었다
|
||||
|
||||
| 버전 | 진실의 원천 | 전체 재시작하면 |
|
||||
|---|---|---|
|
||||
| ~24 | **Infinispan (메모리)** | **세션 전부 소멸** |
|
||||
| 25 | 선택 (`persistent-user-sessions` 옵션) | 설정에 따라 |
|
||||
| **26 (지금 이 실험대)** | **데이터베이스** | **세션 살아남음** |
|
||||
|
||||
**이게 결정적이다.** 인터넷에 있는 Keycloak 클러스터링 자료 대부분은
|
||||
24 이전 기준이라 **"세션은 Infinispan이 노드끼리 복제한다"** 고 쓰여 있다.
|
||||
26에서는 더 이상 사실이 아니다.
|
||||
|
||||
> 제가 처음에 개념 문서에 "Infinispan = 캐시 + 노드 간 실시간 전파"라고
|
||||
> 써둔 것도 이 옛 모델을 그대로 옮긴 것이었다. 실험 0에서 틀렸음이 드러났다.
|
||||
|
||||
---
|
||||
|
||||
## 6. Infinispan / JGroups / 7800 — 이름들의 정체
|
||||
|
||||
실험 로그에 계속 나오는 이름들이다.
|
||||
|
||||
```
|
||||
Keycloak 프로세스
|
||||
┌────────────────────────────────────────┐
|
||||
│ Infinispan "세션을 어디 두고 어떻게 │ ← 캐시 계층
|
||||
│ 나눌까" │
|
||||
│ │ │
|
||||
│ JGroups "누가 우리 멤버이고 │ ← 그룹 통신 계층
|
||||
│ 어떻게 메시지를 주고받나" │
|
||||
│ │ │
|
||||
│ TCP 7800 실제 소켓 │ ← 네트워크
|
||||
└────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
| 이름 | 정체 |
|
||||
|---|---|
|
||||
| **Infinispan** | Keycloak 내장 캐시. 세션·realm 설정·로그인 실패 횟수 등을 담는다 |
|
||||
| **JGroups** | Infinispan이 노드끼리 대화할 때 쓰는 하부 라이브러리 |
|
||||
| **TCP 7800** | JGroups가 쓰는 포트. **노드 간 통신 경로** |
|
||||
| **jdbc-ping** | 서로를 **찾는** 방법. DB의 `JGROUPS_PING` 테이블을 게시판처럼 쓴다 |
|
||||
| `ISPN000094` | "새 멤버 명단을 받았다"는 로그 코드 |
|
||||
|
||||
**찾는 것과 대화하는 것이 다른 경로다.**
|
||||
|
||||
```
|
||||
디스커버리 (서로를 찾는다) → PostgreSQL JGROUPS_PING 테이블
|
||||
트랜스포트 (실제 대화) → TCP 7800
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. 왜 쿠버네티스와 노드 2대가 나오는가
|
||||
|
||||
당신 질문은 "Keycloak이 여러 개일 경우"였다. 그걸 **진짜로** 재현하려면
|
||||
Keycloak 프로세스 2개가 **서로 다른 기계**에 있어야 한다.
|
||||
|
||||
| 방식 | 노드 상실을 실험할 수 있나 |
|
||||
|---|---|
|
||||
| Docker 컨테이너 2개 (한 기계) | **못 한다.** 커널이 하나라 "기계가 죽는" 상황을 못 만든다 |
|
||||
| **VM 2대 + k3s** | **된다.** 하나를 전원 차단할 수 있다 |
|
||||
|
||||
그래서 이 실험대는 VM 2대(`kc-lab-1`, `kc-lab-2`) 위에 k3s를 올렸다.
|
||||
|
||||
```
|
||||
kc-lab-1 (k3s 서버) kc-lab-2 (k3s 에이전트)
|
||||
├─ keycloak-1 ├─ keycloak-0
|
||||
├─ traefik, coredns └─ postgres
|
||||
└─ prometheus, grafana
|
||||
```
|
||||
|
||||
**`keycloak-0` / `keycloak-1` 은 Keycloak 프로세스**이고,
|
||||
**`kc-lab-1` / `kc-lab-2` 는 그것들이 올라간 기계**다. 이름이 비슷해서
|
||||
헷갈리기 쉬운데 계층이 다르다.
|
||||
|
||||
---
|
||||
|
||||
## 8. 그래서 실험 0은 무엇을 알아내려 한 것인가
|
||||
|
||||
### 8-1. 답해야 할 실무 질문
|
||||
|
||||
> Keycloak을 2대로 늘렸다. **한 대가 죽으면 로그인한 사람들은 어떻게 되나?**
|
||||
> **DB가 죽으면?** **노드 사이 네트워크가 끊기면?**
|
||||
|
||||
이 질문들에 답하려면 **정상일 때 무엇이 어디에 있는지**를 먼저 알아야 한다.
|
||||
그게 없으면 장애를 일으켜도 무엇이 왜 깨졌는지 해석할 수 없다.
|
||||
|
||||
### 8-2. 그래서 실험 0의 질문은 두 개다
|
||||
|
||||
```
|
||||
질문 A. 한 노드에서 만든 세션을 다른 노드가 쓸 수 있는가?
|
||||
↓ 답: 그렇다
|
||||
|
||||
질문 B. 그 공유는 무엇 덕분인가?
|
||||
(a) Infinispan 이 메모리를 복제해서
|
||||
(b) 둘 다 같은 DB 를 봐서
|
||||
```
|
||||
|
||||
### 8-3. **B를 구분해야 하는 이유** — 운영 대응이 정반대다
|
||||
|
||||
| 상황 | (a) 복제라면 | (b) DB라면 |
|
||||
|---|---|---|
|
||||
| 노드 간 7800 끊김 | **세션 공유 깨짐** | **멀쩡** |
|
||||
| DB 죽음 | 한동안 버팀 | **즉시 전면 장애** |
|
||||
| 노드 1대 죽음 | 세션 살아남음 | 세션 살아남음 |
|
||||
| 성능 병목 | 노드 간 네트워크 | **DB, 커넥션 풀** |
|
||||
| 튜닝할 곳 | JGroups 설정 | **DB 인덱스, 커넥션 수** |
|
||||
| 노드를 10대로 늘리면 | **복제 트래픽 폭증** | DB 부하 증가 |
|
||||
|
||||
**같은 증상에 정반대 처방이 나온다.** 그래서 추측이 아니라 측정으로
|
||||
확정해야 했다.
|
||||
|
||||
---
|
||||
|
||||
## 9. 왜 하필 "캐시 엔트리 개수"를 셌는가
|
||||
|
||||
질문 B를 가르는 가장 직접적인 방법이기 때문이다.
|
||||
|
||||
```
|
||||
keycloak-0 에만 로그인을 보낸다
|
||||
│
|
||||
└──▶ 그리고 keycloak-1 의 메모리를 들여다본다
|
||||
|
||||
그 세션이 들어와 있으면 → (a) 복제한 것
|
||||
비어 있으면 → (b) DB 로 공유한 것
|
||||
```
|
||||
|
||||
Keycloak은 자기 캐시에 몇 개가 들었는지를 `/metrics` 로 알려준다.
|
||||
|
||||
```
|
||||
vendor_statistics_approximate_entries_unique{cache="sessions"} 7.0
|
||||
───────────── ───
|
||||
세션 캐시 7개 들어 있다
|
||||
```
|
||||
|
||||
**측정 결과: keycloak-1은 계속 0이었다.** keycloak-0이 14개를 들고 있는
|
||||
동안에도 0. 그리고 keycloak-1에 직접 로그인을 보낸 순간에만 늘었다.
|
||||
|
||||
```
|
||||
단계 k0 k1
|
||||
시작 2 0
|
||||
keycloak-1 에 로그인 5회 2 5 ← k0 안 늘어남
|
||||
keycloak-0 에 로그인 5회 7 5 ← k1 안 늘어남
|
||||
|
||||
PostgreSQL 세션 수: 12 = 7 + 5 ← 캐시 합과 정확히 일치
|
||||
```
|
||||
|
||||
**각 노드는 자기가 처리한 것만 캐시한다. 메모리는 건너가지 않는다.**
|
||||
→ 답은 **(b)**.
|
||||
|
||||
그리고 마지막으로 PostgreSQL 로그를 켜서, keycloak-1이 **실제로 날리는
|
||||
SELECT 문**을 잡았다. 추측이 아니라는 것을 못 박기 위해서다.
|
||||
|
||||
---
|
||||
|
||||
## 10. 이 사실이 당신 운영에 뜻하는 것
|
||||
|
||||
| 알게 된 것 | 실무적 의미 |
|
||||
|---|---|
|
||||
| 세션은 DB에 있다 | **DB가 단일 장애점이다.** HA·백업 계획이 Keycloak 대수보다 중요하다 |
|
||||
| 메모리는 로컬 캐시일 뿐 | Keycloak을 몇 대로 늘려도 **노드 간 트래픽은 안 는다.** 대신 DB 부하가 는다 |
|
||||
| 남의 세션은 캐시 안 함 | **sticky session 은 정확성이 아니라 성능 문제다.** 없어도 동작하지만 DB를 더 친다 |
|
||||
| refresh 마다 DB 읽기+쓰기 | access token 수명을 줄이면 **DB 부하가 그만큼 는다** |
|
||||
| `synchronous_commit OFF` | **DB가 강제 종료되면 직전 수백 ms 갱신이 사라진다** (의도된 설계) |
|
||||
| 낙관적 락 (`VERSION`) | **동시에 refresh 하면 한쪽이 진다.** 클라이언트에 재시도가 필요하다 |
|
||||
|
||||
---
|
||||
|
||||
## 11. 앞으로 할 실험과 각각이 답하는 질문
|
||||
|
||||
| # | 실험 | 답하는 실무 질문 | 예측 |
|
||||
|---|---|---|---|
|
||||
| **A-0** | ✅ 세션 복제 확인 | 정상일 때 세션은 어디 있나 | — (완료) |
|
||||
| **A-1** | TCP 7800 차단 | 노드 간 네트워크가 끊기면? | 세션 공유는 **안 깨짐**. 무효화 전파가 깨질 것 |
|
||||
| **A-2** | DB 정지 | **DB가 죽으면?** | **즉시 전면 장애** |
|
||||
| A-2' | DB 강제 종료 | 복구하면 뭘 잃나 | 직전 수백 ms 세션 갱신 소멸 |
|
||||
| **A-3** | 노드 1대 전원 차단 | **Keycloak 한 대가 죽으면?** | 세션 살아남음 |
|
||||
| **A-4** | volatile 모드 비교 | 옛 방식(24 이전)은 뭐가 다른가 | A-1이 **정반대로** 치명적이 됨 |
|
||||
| B-5 | 동시 refresh 경쟁 | 토큰 갱신이 겹치면? | 한쪽이 낙관적 락에서 짐 |
|
||||
| — | SSO 다중 앱 | **SSO를 붙이면 뭐가 달라지나** | Keycloak 세션 하나가 전 앱을 좌우 |
|
||||
|
||||
**A-1이 특히 중요하다.** 통념("클러스터 포트 막으면 세션 깨짐")과
|
||||
이번 측정("세션은 7800으로 안 다님")이 정면으로 어긋나므로, **둘 중 하나는
|
||||
틀렸다.** 실험이 판정한다.
|
||||
|
||||
---
|
||||
|
||||
## 12. 용어 빠른 참조
|
||||
|
||||
| 용어 | 한 줄 |
|
||||
|---|---|
|
||||
| **세션** | 서버가 "이 사람 로그인했음"을 기억하는 것 |
|
||||
| **SSO 세션** | Keycloak이 가진 세션. 이게 죽으면 전 앱 재로그인 |
|
||||
| **access token** | 짧게 사는 JWT. 서명만으로 검증. 취소 불가 |
|
||||
| **refresh token** | access token을 새로 받는 표. **세션 조회가 필요** |
|
||||
| **sid** | 세션 식별자. JWT·DB·관리 API에서 **같은 문자열** |
|
||||
| **Infinispan** | Keycloak **내장** 캐시 (별도 설치 아님) |
|
||||
| **JGroups** | Infinispan의 노드 간 통신 라이브러리 |
|
||||
| **jdbc-ping** | DB 테이블로 서로를 찾는 방식 |
|
||||
| **7800** | 노드 간 통신 포트 |
|
||||
| **persistent-user-sessions** | 세션을 DB에 저장하는 기능. **KC 26 기본값** |
|
||||
| **`OFFLINE_USER_SESSION`** | 이름과 달리 **온라인 세션도** 여기 있다 (`offline_flag='0'`) |
|
||||
| **낙관적 락** | 읽을 때 버전과 같을 때만 쓰기. 충돌은 사후 검출 |
|
||||
| **kc-lab-1/2** | VM(기계) 이름 |
|
||||
| **keycloak-0/1** | Keycloak 프로세스(파드) 이름 |
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user