The experiment documents record what was found. These record what to type to
reproduce it, in folders per stage.
Two kinds of command are kept apart. 하기/확인 is what somebody actually types
at a terminal — short, one at a time. 근거를 재려면 is the long measuring form
this lab used to put evidence in a document, marked as not needed day to day.
The same split applies to curl: -I to look once, -w '%{http_code}' only when
comparing across repetitions.
No placeholders. Where a value is needed the command that produces it is
given, and secrets are checked by length rather than printed:
TOKEN=$(ssh kc-lab-1 'sudo cat /var/lib/rancher/k3s/server/node-token')
echo "${#TOKEN} 자"
Stage 05 verifies resources in layers, because a Secret existing and a pod
having received it are different facts: keys, then length, then the value
inside the container, then which env var came from which Secret. Same for
workloads — Deployment to ReplicaSet to Pod, with the seven ReplicaSets this
cluster actually carries as the worked example.
Two commands were wrong and re-running them caught it. kubectl get endpoints
prints a deprecation warning on v1.33+, so the guide uses describe svc and
EndpointSlice. And the Keycloak image has no curl, so reading metrics from
inside the container fails with exit 127 — the guide asks Prometheus instead,
or runs a throwaway curl pod.
Read-only checks were executed against the running lab and their output is
quoted verbatim. Creating commands could not be re-run without destroying the
lab, so they are the ones used at build time; the README says which is which.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
132 lines
3.9 KiB
Markdown
132 lines
3.9 KiB
Markdown
# 06 — 관측
|
|
|
|
## 이 단계가 끝나면
|
|
|
|
Prometheus 가 Keycloak 을 긁고, `vendor_cluster_size` 로 클러스터 상태를
|
|
밖에서 볼 수 있다.
|
|
|
|
## 전제
|
|
|
|
[05](../05-keycloak/) 가 끝나 Keycloak 두 노드가 떴다.
|
|
|
|
## 왜 필요한가
|
|
|
|
실험의 판정을 **밖에서만** 하면 놓친다. A-1 에서 7800 을 끊었는데 외부 응답이
|
|
전부 200 이었다 — 분단된 노드가 스스로 로드밸런서에서 빠졌기 때문이다.
|
|
클러스터 안을 보는 눈이 따로 있어야 한다.
|
|
|
|
---
|
|
|
|
## 1. 적용
|
|
|
|
**하기**
|
|
```bash
|
|
kubectl apply -f deploy/lab/k8s/observability.yaml
|
|
kubectl -n observability rollout status deploy/prometheus --timeout=180s
|
|
```
|
|
|
|
**확인**
|
|
```bash
|
|
kubectl -n observability get pods
|
|
```
|
|
|
|
**실측**
|
|
```
|
|
grafana-845b5678cf-b6gvc 1/1 Running
|
|
node-exporter-9qk9w 1/1 Running
|
|
node-exporter-c2mz4 1/1 Running
|
|
prometheus-6774f94f7c-pzr2t 1/1 Running
|
|
```
|
|
|
|
node-exporter 가 **둘**인 것은 DaemonSet 이라 노드마다 하나씩 뜨기 때문이다.
|
|
|
|
## 2. 무엇을 긁고 있나 — 여기가 중요하다
|
|
|
|
**확인**
|
|
```bash
|
|
kubectl -n observability exec deploy/prometheus -- \
|
|
wget -qO- localhost:9090/api/v1/targets | grep -o '"job":"[^"]*"' | sort -u
|
|
```
|
|
|
|
**실측**
|
|
```
|
|
"job":"keycloak"
|
|
"job":"kubelet"
|
|
"job":"node-exporter"
|
|
"job":"prometheus"
|
|
```
|
|
|
|
**★ Redis · BFF · PostgreSQL 이 없다.** 이 실험대는 그것들을 긁지 않는다.
|
|
그래서 B층 실험 대부분에 Grafana 화면이 없는데, **안 찍은 것이 아니라 지표가
|
|
없는 것**이다.
|
|
|
|
> 이것을 「스크린샷 누락」이 아니라 **측정된 공백**으로 기록했다.
|
|
> [`evidence/followup/04-observability-gap.txt`](../../evidence/followup/04-observability-gap.txt)
|
|
|
|
## 3. 클러스터 상태를 본다
|
|
|
|
**하기**
|
|
```bash
|
|
kubectl -n observability exec deploy/prometheus -- \
|
|
wget -qO- 'localhost:9090/api/v1/query?query=vendor_cluster_size'
|
|
```
|
|
|
|
**실측**
|
|
```
|
|
keycloak-1 → 2
|
|
keycloak-0 → 2
|
|
```
|
|
|
|
**두 노드가 각각 자기가 아는 멤버 수를 보고한다.** 분단되면 한쪽은 2,
|
|
다른 쪽은 1 이 된다 — **한 노드만 보면 분단을 놓친다.**
|
|
|
|
자주 보는 지표들이다.
|
|
|
|
| 지표 | 무엇 |
|
|
|---|---|
|
|
| `vendor_cluster_size` | 이 노드가 아는 멤버 수 |
|
|
| `vendor_jgroups_*` | JGroups 프로토콜별 카운터 |
|
|
| `vendor_statistics_approximate_entries_unique{cache="sessions"}` | 이 노드의 세션 캐시 엔트리 수 |
|
|
| `agroal_*` | JDBC 커넥션 풀 |
|
|
| `up` | 스크레이프 성공 여부 |
|
|
|
|
## 4. `up` 을 믿지 않는다
|
|
|
|
**A-2 에서 503 이 나는 동안에도 `up` 은 1 이었다.** 프로세스가 살아 있고
|
|
`/metrics` 가 응답하기만 하면 1 이므로 **「살아 있지만 쓸모없는」 상태를
|
|
보지 못한다.**
|
|
|
|
```bash
|
|
kubectl -n observability exec deploy/prometheus -- \
|
|
wget -qO- 'localhost:9090/api/v1/query?query=up'
|
|
```
|
|
|
|
경보를 걸 때는 `up == 0` 만으로 부족하고 **기능 지표**를 함께 본다.
|
|
|
|
```bash
|
|
curl -s -o /dev/null -w '%{http_code}\n' https://auth.hyeonworks.com/realms/master
|
|
```
|
|
|
|
## 5. Grafana 를 볼 때
|
|
|
|
**하기** — 밖에 열지 않고 포트포워드로 본다
|
|
```bash
|
|
kubectl -n observability port-forward svc/grafana 3000:3000
|
|
```
|
|
|
|
브라우저에서 `http://localhost:3000`.
|
|
|
|
> 실험 중에는 Grafana 보다 **Prometheus 쿼리 API** 가 편하다. 값을 그대로
|
|
> 뽑아 비교할 수 있고 스크린샷보다 근거로 남기기 좋다.
|
|
|
|
---
|
|
|
|
## 막히면
|
|
|
|
| 증상 | 원인 | 확인 |
|
|
|---|---|---|
|
|
| Keycloak 지표가 안 보임 | 9000 이 안 열렸거나 스크레이프 설정 누락 | 위 2번 targets |
|
|
| 값이 한 노드만 나옴 | 다른 노드 스크레이프 실패 | targets 의 `health` 필드 |
|
|
| 컨테이너 안에서 curl 실패 | **Keycloak 이미지에 curl 이 없다** | 밖에서 Prometheus 로 묻는다 |
|
|
| Grafana 에 데이터 없음 | 데이터소스 주소 오류 | Prometheus 서비스 이름 확인 |
|