18 KiB
18 KiB
operations / runbook / upgrade / rollback 예시
좋은 예시 1: 표준 application 변경 절차
# 1) render
kubectl kustomize gitops/clusters/prod/main/stages/50-apps > /tmp/render.yaml
# 2) diff
kubectl diff -k gitops/clusters/prod/main/stages/50-apps
# 3) apply
kubectl apply -k gitops/clusters/prod/main/stages/50-apps
# 4) rollout status with timeout
kubectl rollout status deployment/auth-server -n auth-prod --timeout=10m
# 5) smoke test
curl -fsS https://auth.internal.example.com/actuator/health/readiness
# 6) SLO dashboard check (p99 latency, error rate)
왜 좋은가:
- render / diff / apply / status / post-check 가 명시적으로 분리.
--timeout으로 무한 대기 방지.- post-check가 단순 curl 이 아니라 readiness endpoint 대상.
좋은 예시 2: Deployment rollingUpdate 파라미터 워크로드별 튜닝
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: auth-server
namespace: auth-prod
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
spec:
replicas: 10
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25%
maxUnavailable: 10% # latency-sensitive면 0
selector:
matchLabels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
template:
metadata:
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: auth-server
image: registry.example.com/identity/auth-server:1.25.0
ports:
- { name: http, containerPort: 8080 }
resources:
requests: { cpu: 500m, memory: 1Gi }
limits: { memory: 1536Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ["ALL"] }
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: legacy-singleton
namespace: legacy
spec:
replicas: 1
strategy:
type: Recreate # singleton이며 동시성 금지
selector:
matchLabels:
app.kubernetes.io/name: legacy-singleton
template:
metadata:
labels:
app.kubernetes.io/name: legacy-singleton
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: app
image: registry.example.com/legacy/singleton:1.0.0
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 512Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ["ALL"] }
왜 좋은가:
- fleet 규모에 맞춘 maxSurge/maxUnavailable.
- singleton 에 Recreate (PVC ReadWriteOnce 전제 충족).
좋은 예시 3: Argo Rollouts canary with AnalysisTemplate
---
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: success-rate
namespace: auth-prod
spec:
args:
- name: service-name
metrics:
- name: success-rate
interval: 1m
count: 5
successCondition: result[0] >= 0.99
failureLimit: 2
provider:
prometheus:
address: http://prometheus.monitoring.svc:9090
query: |
sum(rate(http_requests_total{service="{{args.service-name}}",status_class=~"2.."}[2m]))
/
sum(rate(http_requests_total{service="{{args.service-name}}"}[2m]))
- name: p99-latency
interval: 1m
count: 5
successCondition: result[0] <= 0.5
failureLimit: 2
provider:
prometheus:
address: http://prometheus.monitoring.svc:9090
query: |
histogram_quantile(0.99,
sum by (le) (rate(http_request_duration_seconds_bucket{service="{{args.service-name}}"}[2m]))
)
---
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: auth-server
namespace: auth-prod
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
spec:
replicas: 10
revisionHistoryLimit: 5
selector:
matchLabels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
template:
metadata:
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
spec:
containers:
- name: auth-server
image: registry.example.com/identity/auth-server:1.25.0
ports:
- name: http
containerPort: 8080
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
memory: "1536Mi"
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: http
strategy:
canary:
canaryService: auth-server-canary
stableService: auth-server-stable
trafficRouting:
nginx:
stableIngress: auth-server
steps:
- setWeight: 10
- pause: { duration: 2m }
- analysis:
templates:
- templateName: success-rate
args:
- name: service-name
value: auth-server
- setWeight: 25
- pause: { duration: 5m }
- analysis:
templates:
- templateName: success-rate
args:
- name: service-name
value: auth-server
- setWeight: 50
- pause: { duration: 10m }
- analysis:
templates:
- templateName: success-rate
args:
- name: service-name
value: auth-server
- setWeight: 100
---
apiVersion: v1
kind: Service
metadata:
name: auth-server-stable
namespace: auth-prod
spec:
selector:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
ports:
- name: http
port: 80
targetPort: http
---
apiVersion: v1
kind: Service
metadata:
name: auth-server-canary
namespace: auth-prod
spec:
selector:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
ports:
- name: http
port: 80
targetPort: http
왜 좋은가:
AnalysisTemplate이 Prometheus success-rate + p99 latency 를 동시에 측정.failureLimit: 2→ 두 번 실패 시 자동 abort.- canary step: 10% → 25% → 50% → 100% 각 단계에 pause + analysis.
- stable/canary Service 두 개 + NGINX ingress traffic routing.
좋은 예시 4: K3s System Upgrade Controller Plan (server + agent)
---
apiVersion: v1
kind: Namespace
metadata:
name: system-upgrade
---
apiVersion: v1
kind: Secret
metadata:
name: k3s-upgrade-token
namespace: system-upgrade
type: Opaque
stringData:
# 실제 환경은 K3S_TOKEN 값
token: "REPLACE_WITH_NODE_TOKEN"
---
apiVersion: upgrade.cattle.io/v1
kind: Plan
metadata:
name: k3s-server
namespace: system-upgrade
labels:
k3s-upgrade: server
spec:
concurrency: 1
nodeSelector:
matchExpressions:
- { key: node-role.kubernetes.io/control-plane, operator: In, values: ["true"] }
serviceAccountName: system-upgrade
cordon: true
drain:
force: true
deleteEmptydirData: true
ignoreDaemonSets: true
skipWaitForDeleteTimeout: 60
upgrade:
image: rancher/k3s-upgrade
version: v1.30.3+k3s1
---
apiVersion: upgrade.cattle.io/v1
kind: Plan
metadata:
name: k3s-agent
namespace: system-upgrade
labels:
k3s-upgrade: agent
spec:
concurrency: 1
nodeSelector:
matchExpressions:
- { key: node-role.kubernetes.io/control-plane, operator: NotIn, values: ["true"] }
serviceAccountName: system-upgrade
prepare:
image: rancher/k3s-upgrade
args: ["prepare", "k3s-server"] # server plan 완료 대기
cordon: true
drain:
force: true
deleteEmptydirData: true
ignoreDaemonSets: true
skipWaitForDeleteTimeout: 60
upgrade:
image: rancher/k3s-upgrade
version: v1.30.3+k3s1
왜 좋은가:
- server-plan → agent-plan 분리 + agent 가
prepare로 server 완료 대기. concurrency: 1→ 한 번에 한 노드만 업그레이드 (가용성 보호).cordon + drain→ PDB 존중.deleteEmptydirData: true, ignoreDaemonsets: true표준.version명시 (channel 사용 시 의도치 않은 upgrade 가능).
좋은 예시 5: ArgoCD sync wave + PreSync migration hook
---
apiVersion: batch/v1
kind: Job
metadata:
name: flyway-migrate
namespace: auth-prod
annotations:
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
argocd.argoproj.io/sync-wave: "-1"
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
app.kubernetes.io/component: db-migration
spec:
backoffLimit: 0
activeDeadlineSeconds: 600
ttlSecondsAfterFinished: 86400
template:
spec:
restartPolicy: Never
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
fsGroup: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: flyway
image: flyway/flyway:10.15.0
args: ["-url=jdbc:postgresql://postgres:5432/auth", "validate", "info", "migrate"]
envFrom:
- secretRef:
name: auth-db
resources:
requests: { cpu: 100m, memory: 128Mi }
limits: { memory: 512Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
volumeMounts:
- { name: tmp, mountPath: /tmp }
volumes:
- name: tmp
emptyDir: {}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: auth-server
namespace: auth-prod
annotations:
argocd.argoproj.io/sync-wave: "0"
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
# ... (생략)
---
apiVersion: batch/v1
kind: Job
metadata:
name: smoke-test
namespace: auth-prod
annotations:
argocd.argoproj.io/hook: PostSync
argocd.argoproj.io/hook-delete-policy: HookSucceeded
argocd.argoproj.io/sync-wave: "1"
spec:
backoffLimit: 2
activeDeadlineSeconds: 300
ttlSecondsAfterFinished: 3600
template:
spec:
restartPolicy: OnFailure
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
fsGroup: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: smoke
image: registry.example.com/tools/smoke:1.4.0
args: ["--target", "https://auth.internal.example.com"]
resources:
requests: { cpu: 50m, memory: 64Mi }
limits: { memory: 128Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
왜 좋은가:
- PreSync Job 으로 Flyway migrate 가 app rollout 앞 단계에 실행.
- PostSync Job 으로 smoke test 자동 실행.
- sync-wave 로 순서 명시 (-1 → 0 → 1).
BeforeHookCreation으로 이전 Job 충돌 방지.
좋은 예시 6: blue/green via two Services (수동 패턴)
---
apiVersion: v1
kind: Service
metadata:
name: auth-server # live traffic
namespace: auth-prod
spec:
selector:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
version: blue # <- 이 label만 바꾸면 cutover
ports:
- name: http
port: 80
targetPort: http
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: auth-server-blue
namespace: auth-prod
spec:
replicas: 3
selector:
matchLabels:
app.kubernetes.io/name: auth-server
version: blue
template:
metadata:
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
version: blue
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: auth-server
image: registry.example.com/identity/auth-server:1.24.0
resources:
requests: { cpu: 500m, memory: 1Gi }
limits: { memory: 1536Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: auth-server-green
namespace: auth-prod
spec:
replicas: 3
selector:
matchLabels:
app.kubernetes.io/name: auth-server
version: green
template:
metadata:
labels:
app.kubernetes.io/name: auth-server
app.kubernetes.io/instance: auth-server-prod
version: green
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: auth-server
image: registry.example.com/identity/auth-server:1.25.0
resources:
requests: { cpu: 500m, memory: 1Gi }
limits: { memory: 1536Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
Cutover:
kubectl patch svc auth-server -n auth-prod \
-p '{"spec":{"selector":{"app.kubernetes.io/name":"auth-server","app.kubernetes.io/instance":"auth-server-prod","version":"green"}}}'
왜 좋은가:
- Service selector version label 하나로 전환 / rollback.
- canary 가 아니라 instant cutover.
- 데이터 호환성이 깨진 경우만 사용.
좋은 예시 7: node maintenance flow
NODE=worker-3
# 1) cordon
kubectl cordon "${NODE}"
# 2) drain (PDB 존중)
kubectl drain "${NODE}" \
--ignore-daemonsets \
--delete-emptydir-data \
--grace-period=30 \
--timeout=10m
# 3) 작업 수행 (OS patch, reboot, ...)
# 4) 복귀
kubectl uncordon "${NODE}"
# 5) 재배치 확인
kubectl get pods -A -o wide --field-selector spec.nodeName="${NODE}"
왜 좋은가:
- cordon → drain → uncordon 표준 시퀀스.
- PDB 위반 시 drain 이 대기,
--timeout=10m로 무한 대기 방지. - 플래그 조합이 표준.
좋은 예시 8: Git revision rollback
# 1) 이전 release tag 체크아웃
git checkout v1.24.0
# 2) diff
kubectl diff -k gitops/clusters/prod/main/stages/50-apps
# 3) apply
kubectl apply -k gitops/clusters/prod/main/stages/50-apps
# 4) rollout status
kubectl rollout status deployment/auth-server -n auth-prod --timeout=10m
왜 좋은가:
- live-cluster 수정이 아니라 declarative source of truth 기준.
- 재현 가능.
kubectl rollout undo대비 audit trail 이 명확 (Git commit 기반).
나쁜 예시 1: diff 없이 apply
kubectl apply -k gitops/clusters/prod/main/stages/50-apps
문제:
- 실제 변경 범위를 모른 채 적용.
- review / 승인 / 검증 프로세스 약화.
- 의도치 않은 리소스 삭제/수정 가능 (특히 pruned resource).
Fix: kubectl diff -k 선행.
나쁜 예시 2: migration을 app startup에 숨김
# Deployment container
command: ["/bin/sh", "-c", "flyway migrate && java -jar app.jar"]
문제:
- app rollout 실패와 schema 변경 실패가 섞임.
- rollout 중 여러 replica 가 동시에 migrate → race condition / lock contention.
- 롤백 시 schema 변경이 남음.
Fix: PreSync Job 또는 별도 CI 단계로 Flyway migrate 를 분리.
나쁜 예시 3: rollout undo 로 DB rollback 기대
kubectl rollout undo deployment/auth-server
# ... 이제 DB schema 도 되돌아갔을 것이다?
문제:
- rollout undo 는 workload pod template 만 되돌린다.
- schema 변경은 남아 있음 → 이전 버전 app이 새 schema 와 mismatch → 500 error.
- rollback ≠ DB rollback.
Fix: schema 는 expand/contract 패턴으로 forward-compatible. 이전 버전 코드가 새 schema 에서도 동작하도록 릴리스를 분리.
나쁜 예시 4: 운영 노드 manifests 디렉터리 직접 편집
ssh k3s-server-1
vim /var/lib/rancher/k3s/server/manifests/auth-server.yaml
문제:
- Git source of truth 우회.
- 멀티 서버 간 동기화 없음.
- packaged AddOn 동작과 충돌 가능.
- ArgoCD 가 drift 로 인식하고 되돌릴 수 있음.
Fix: Git PR → render → diff → apply 흐름.
나쁜 예시 5: Recreate strategy 를 stateless app에 사용
apiVersion: apps/v1
kind: Deployment
metadata:
name: auth-server
spec:
replicas: 5
strategy:
type: Recreate # BAD - stateless 인데 downtime 발생
문제:
- 모든 replica 동시 종료 → full downtime.
- rolling update 의 장점 (점진 전환, rollback 용이) 상실.
Fix: stateless app은 RollingUpdate + 워크로드별 maxSurge/maxUnavailable 튜닝.
나쁜 예시 6: PDB 없이 drain
kubectl drain worker-3 --ignore-daemonsets --delete-emptydir-data
문제:
- PDB 가 없으면 critical workload 가 동시에 evict → downtime.
- 특히 replica < 3 이면 완전 손실.
Fix: PDB 설계 선결 조건. 좋은 예시 7 참조.
나쁜 예시 7: kubectl rollout status 에 timeout 없음
kubectl rollout status deployment/auth-server -n auth-prod
# 무한 대기 가능
문제:
- rollout 이 hang 상태일 때 CI/CD pipeline 이 무한 대기.
- 자동화 실패 원인이 숨는다.
Fix: 항상 --timeout=10m (워크로드별 조정).