본문으로 건너뛰기
AIDevOps
  • Learn
  • Learning Paths
  • Practice
  • Open Source
  • Books
  • Engineering

    AI DevOpsAI 서비스 개발·운영 전체 지도LLMOpsLLM 배포·평가·관측실전 프로젝트AI Agent 프로젝트 실습

    Knowledge

    Docs기술 문서 모음Blog엔지니어링 아티클Plogger개발 기록 피드

    Validate

    Certification3단계 역량 인증 · 준비 중
AI Models
LlamaMistralGemmaDeepSeekQwen
🐳 DevOps
DevOps 입문 & 로드맵LinuxDockerCI/CD|Kubernetes 기본K8s 심화/실무PrometheusGrafana
🤖 AI 실전 개발
AI 실전 입문 & 로드맵Hugging FaceLangChainLlamaIndexLLMOps|LangGraphMCPMulti-AgentAgent Evaluation
🧠 AI Core
AI 입문 & 로드맵ML FundamentalsLLM Fundamentals|Python AIC++|PyTorchTensorFlowJAX
🧠 AI Agent 개발
금융 AI AgentLLM API 서버주식 투자 AgentAIOps AI Agent교육 AI Agent코딩 AI Agent
🌱 Spring Cloud
Spring 입문 & 로드맵Spring Cloud GatewaySpring BootJava|Spring AISpring SecuritySpring BatchSpring JPA
🧱 인프라
인프라 입문 & 로드맵NginxRedis
☁️ 클라우드
클라우드 입문 & 로드맵AWSGCPAzureNCPCloudflare
🎨 Frontend
Frontend 입문 & 로드맵JavaScriptTypeScript|ReactNext.js|VueNuxt
📱 Mobile
Mobile 입문 & 로드맵KotlinAndroidFlutter
⚙️ Backend
Backend 입문 & 로드맵Python 기본FastAPIDjangoFlask|CGoGinNode.js
💾 Database
DB 입문 & 로드맵공통 SQLOracleMySQLPostgreSQL|MongoDB벡터 DB
🧪 검증
k6JMeternGrinder
AIDevOps

Engineering AI. From Code to Production.
AI와 AI Agent를 개발하고 운영하기 위한 엔지니어링 학습 플랫폼

Learn

  • 전체 가이드
  • Learning Paths
  • Practice
  • Books

Resources

  • AI DevOps
  • LLMOps
  • 실전 프로젝트
  • Docs
  • Blog
  • Plogger
  • Open Source
  • Certification (준비 중)

Start Here

  • AI Core 로드맵
  • AI 실전 개발 로드맵
  • Spring Cloud 로드맵
  • DevOps 로드맵
  • 인프라 로드맵

 

  • 클라우드 로드맵
  • Frontend 로드맵
  • Mobile 로드맵
  • Backend 로드맵
  • Database 로드맵
© 2026 AI DevOps Korea. All rights reserved.
이용약관개인정보처리방침Sitemaptestforge.kr
  1. Home
  2. Learn
  3. DevOps
  4. Kubernetes 심화/실무
프로덕션 운영 & 심화 가이드

☸️ Kubernetes 심화/실무 완전 가이드

Visitors

HPA/VPA 오토스케일링, RBAC 보안, StatefulSet·DaemonSet·CronJob 워크로드, GPU 기반 LLM 추론 서버 배포와 카나리 롤아웃·큐 기반 오토스케일링 같은 AI 모델 운영 전략, 업그레이드·장애 복구, 차트 구조·Hook·서브차트까지 다루는 Helm 심화, Contour/Gateway API 기반 Ingress 전환, LB·Contour 버전별 Proxy Protocol 설정과 내부 호출 대응, 모니터링까지 — 프로덕션 Kubernetes 운영에 필요한 심화 주제를 다룹니다.

  • Advanced · 심화
  • 업데이트 2026.09.20
  • 약 107분 읽기
  • 22개 섹션
  • 예제 코드 77개

포함된 Learning Path

이 가이드는 아래 경로의 한 단계입니다. 앞뒤 순서와 함께 학습해보세요.

  • Cloud Native Engineer →
  • AI Platform Engineer →
자동 스케일링 & 무중단 배포RBAC 기반 권한 관리StatefulSet/DaemonSet/CronJob 운영GPU 노드 기반 LLM 추론 서버 & AI 모델 운영 전략Helm 차트 심화 & GitOps 연계Contour/Gateway API 기반 Ingress 전환 & 정책 설정Proxy Protocol 기반 클라이언트 IP 보존 & 내부 호출 대응

관련 프레임워크 & 개발환경

☸️Kubernetes 기본→🐳Docker→CICI/CD→📊Prometheus→📈Grafana→

목차

0 / 24
  1. 가이드 사용법
  2. 구조 다이어그램
  3. HPA & VPA 오토스케일링
  4. RBAC & 보안
  5. StatefulSet vs Stateless
  6. StatefulSet 완전 가이드
  7. StatefulSet 배포 워크플로
  8. 실전 샘플
  9. StatefulSet 심화 — Parallel · 볼륨 확장 · 스냅샷
  10. DaemonSet 완전 가이드
  11. CronJob & Job 완전 가이드
  12. Java Agent StatefulSet 배포
  13. GPU 워크로드 & LLM 추론 서버 배포
  14. AI 모델 배포 전략 & 운영
  15. 운영 & 업그레이드
  16. Helm 차트
  17. Contour & Gateway API — NGINX Ingress 대체
  18. Proxy Protocol — 클라이언트 IP 보존
  19. LB별 Proxy Protocol 설정
  20. Contour 버전별 Envoy 설정 (1.28~1.33)
  21. PROXY 헤더 없는 내부 호출 & 대체 방법
  22. Proxy Protocol 검증 & 트러블슈팅
  23. Contour/Envoy 모니터링 — Prometheus & Grafana
  24. 다음 단계
목차 24개 섹션
  1. 가이드 사용법
  2. 구조 다이어그램
  3. HPA & VPA 오토스케일링
  4. RBAC & 보안
  5. StatefulSet vs Stateless
  6. StatefulSet 완전 가이드
  7. StatefulSet 배포 워크플로
  8. 실전 샘플
  9. StatefulSet 심화 — Parallel · 볼륨 확장 · 스냅샷
  10. DaemonSet 완전 가이드
  11. CronJob & Job 완전 가이드
  12. Java Agent StatefulSet 배포
  13. GPU 워크로드 & LLM 추론 서버 배포
  14. AI 모델 배포 전략 & 운영
  15. 운영 & 업그레이드
  16. Helm 차트
  17. Contour & Gateway API — NGINX Ingress 대체
  18. Proxy Protocol — 클라이언트 IP 보존
  19. LB별 Proxy Protocol 설정
  20. Contour 버전별 Envoy 설정 (1.28~1.33)
  21. PROXY 헤더 없는 내부 호출 & 대체 방법
  22. Proxy Protocol 검증 & 트러블슈팅
  23. Contour/Envoy 모니터링 — Prometheus & Grafana
  24. 다음 단계

가이드 사용법

읽는 방향

Kubernetes 심화/실무를 실무 흐름으로 이해하기

HPA/VPA 오토스케일링, RBAC 보안, StatefulSet·DaemonSet·CronJob 워크로드, GPU 기반 LLM 추론 서버 배포와 카나리 롤아웃·큐 기반 오토스케일링 같은 AI 모델 운영 전략, 업그레이드·장애 복구, 차트 구조·Hook·서브차트까지 다루는 Helm 심화, Contour/Gateway API 기반 Ingress 전환, LB·Contour 버전별 Proxy Protocol 설정과 내부 호출 대응, 모니터링까지 — 프로덕션 Kubernetes 운영에 필요한 심화 주제를 다룹니다. 이 가이드는 개념을 나열하기보다, 실제 프로젝트에서 판단해야 하는 순서대로 내용을 따라갈 수 있게 구성했습니다.

핵심 관점

인프라 / 운영

설치 명령을 외우기보다 트래픽, 런타임, 관측, 장애 대응이 어떤 순서로 이어지는지 파악합니다.

자동 스케일링 & 무중단 배포RBAC 기반 권한 관리StatefulSet/DaemonSet/CronJob 운영GPU 노드 기반 LLM 추론 서버 & AI 모델 운영 전략Helm 차트 심화 & GitOps 연계Contour/Gateway API 기반 Ingress 전환 & 정책 설정Proxy Protocol 기반 클라이언트 IP 보존 & 내부 호출 대응

구조 다이어그램

글로 읽은 내용을 머릿속에 오래 남기려면 먼저 흐름을 그림으로 잡는 편이 좋습니다. 아래 두 그림은 Kubernetes 심화/실무를 학습할 때 계속 되돌아볼 수 있는 기준 지도입니다.

학습 흐름

다이어그램 렌더링 중…

아키텍처 관점

다이어그램 렌더링 중…

HPA & VPA 오토스케일링

Kubernetes 심화/실무를 처음 펼칠 때는 세부 명령보다 큰 그림이 먼저입니다. 이 섹션에서는 앞으로 배울 개념들이 어떤 문제를 풀기 위해 등장했는지부터 잡아봅니다.

트래픽이 일정하지 않은 서비스는 고정 replicas 대신 오토스케일러로 부하에 맞춰 Pod 수(HPA)나 컨테이너 리소스(VPA)를 자동 조정해야 합니다. 두 방식을 같은 리소스에 CPU/메모리 타깃으로 동시에 켜면 서로 충돌하므로, HPA는 CPU/커스텀 메트릭 기준으로, VPA는 요청값 추천 전용(updateMode: Off)으로 분리 운용하는 것이 안전합니다.
다이어그램 렌더링 중…
hpa.yamlYAML
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: myapp-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: myapp
  minReplicas: 2
  maxReplicas: 20
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60        # CPU 60% 초과 시 스케일 아웃
  - type: Resource
    resource:
      name: memory
      target:
        type: AverageValue
        averageValue: 400Mi
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300 # 5분 관찰 후 축소 — 스케일 플래핑 방지
      policies:
      - type: Pods
        value: 1
        periodSeconds: 60
    scaleUp:
      stabilizationWindowSeconds: 0
      policies:
      - type: Percent
        value: 100
        periodSeconds: 15
BASH
# Metrics Server 설치 (HPA 필수 의존성)
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

# 상태 확인
kubectl get hpa myapp-hpa --watch
kubectl describe hpa myapp-hpa          # 현재 메트릭 값, 스케일 이벤트 확인
kubectl top pods                        # Pod별 실시간 CPU/메모리
vpa.yamlYAML
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: myapp-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: myapp
  updatePolicy:
    updateMode: "Off"   # 추천값만 계산, 실제 적용은 수동 — 운영 안전
  resourcePolicy:
    containerPolicies:
    - containerName: '*'
      minAllowed: { cpu: 50m, memory: 64Mi }
      maxAllowed: { cpu: 2, memory: 4Gi }
오토스케일러조정 대상전제 조건주의점
HPAPod 복제본 수 (가로 확장)Metrics Server 또는 Prometheus Adapter커스텀 메트릭 사용 시 스케일 다운 지연(stabilizationWindow) 설정 필요
VPA컨테이너 requests/limits (세로 확장)vertical-pod-autoscaler 컴포넌트updateMode: Auto는 Pod 재시작을 유발 — 상태 저장 워크로드엔 위험
Cluster Autoscaler워커 노드 수클라우드 Provider Auto Scaling Group 연동Pod가 스케줄 불가(Pending) 상태여야 노드 증설이 트리거됨

RBAC & 보안

여기서는 RBAC & 보안을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Role-Based Access Control(RBAC)로 팀별, 서비스별 API 접근 권한을 분리합니다. ServiceAccount는 Pod 내부 프로세스 권한 제어에, Role/ClusterRole은 사용자·팀 권한 분리에 사용합니다.
team-rbac.yamlYAML
# 팀별 RBAC — 개발팀은 production namespace 읽기만 허용
apiVersion: v1
kind: ServiceAccount
metadata:
  name: dev-team-reader
  namespace: production
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: read-only
  namespace: production
rules:
- apiGroups: [""]
  resources: ["pods", "services", "endpoints", "configmaps"]
  verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
  resources: ["deployments", "replicasets", "statefulsets"]
  verbs: ["get", "list", "watch"]
- apiGroups: [""]
  resources: ["pods/log"]
  verbs: ["get", "list"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: dev-team-binding
  namespace: production
subjects:
- kind: ServiceAccount
  name: dev-team-reader
  namespace: production
roleRef:
  kind: Role
  name: read-only
  apiGroup: rbac.authorization.k8s.io
ops-clusterrole.yamlYAML
# 운영팀 ClusterRole — 전체 네임스페이스 관리 권한
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ops-admin
rules:
- apiGroups: ["*"]
  resources: ["*"]
  verbs: ["*"]
- nonResourceURLs: ["*"]
  verbs: ["*"]
---
# CI/CD 시스템 전용 ServiceAccount (Deployment만 업데이트)
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: cicd-deployer
rules:
- apiGroups: ["apps"]
  resources: ["deployments", "statefulsets"]
  verbs: ["get", "list", "patch", "update"]
- apiGroups: [""]
  resources: ["services", "configmaps"]
  verbs: ["get", "list", "create", "update", "patch"]
---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: cicd-deployer
  namespace: production
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: cicd-deployer-binding
subjects:
- kind: ServiceAccount
  name: cicd-deployer
  namespace: production
roleRef:
  kind: ClusterRole
  name: cicd-deployer
  apiGroup: rbac.authorization.k8s.io
kubeconfig-user.shBASH
# 외부 사용자(개발자)용 kubeconfig 생성
# 1. 개인키 & CSR 생성
openssl genrsa -out devuser.key 4096
openssl req -new -key devuser.key -out devuser.csr \
  -subj "/CN=devuser/O=dev-team"

# 2. K8s CertificateSigningRequest 제출
cat > devuser-csr.yaml << EOF
apiVersion: certificates.k8s.io/v1
kind: CertificateSigningRequest
metadata:
  name: devuser
spec:
  request: $(base64 -w 0 devuser.csr)
  signerName: kubernetes.io/kube-apiserver-client
  expirationSeconds: 7776000   # 90일
  usages:
  - client auth
EOF
kubectl apply -f devuser-csr.yaml

# 3. CSR 승인
kubectl certificate approve devuser

# 4. 인증서 추출 & kubeconfig 생성
kubectl get csr devuser -o jsonpath='{.status.certificate}' | base64 -d > devuser.crt

kubectl config set-credentials devuser \
  --client-key=devuser.key \
  --client-certificate=devuser.crt \
  --embed-certs=true

kubectl config set-context devuser-context \
  --cluster=kubernetes \
  --user=devuser \
  --namespace=production

# 5. RoleBinding으로 권한 부여
kubectl create rolebinding devuser-read \
  --role=read-only \
  --user=devuser \
  --namespace=production

StatefulSet vs Stateless(Deployment) 비교

여기서는 StatefulSet vs Stateless(Deployment) 비교을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Kubernetes에서 워크로드를 선택할 때 가장 중요한 기준은 "Pod이 상태(데이터)를 직접 보관하는가"입니다. Deployment는 Pod을 교체 가능한 복제본으로 다루지만, StatefulSet은 각 Pod에 고유한 ID, 안정적 네트워크 이름, 독립적 스토리지를 부여합니다.
다이어그램 렌더링 중…
구분Deployment (Stateless)StatefulSet (Stateful)
Pod 이름랜덤 suffix (app-7d4f9c-xkz)순번 고정 (app-0, app-1, app-2)
Pod 교체랜덤 순서로 생성·삭제 가능순서 보장 (0→1→2 생성, 2→1→0 삭제)
네트워크 IDService ClusterIP 공유, Pod IP 가변Headless Service로 app-0.svc, app-1.svc 고정 DNS
스토리지Pod 삭제 시 데이터 소멸 (PVC 미보장)volumeClaimTemplates → 각 Pod 전용 PVC 유지
스케일 아웃순서 없이 즉시 병렬 확장0, 1, 2 순서로 순차 확장
롤링 업데이트랜덤 교체 (maxSurge/maxUnavailable)역순 (2→1→0) 순차 교체
주요 용도API 서버, 웹 앱, AI 추론 서버 등DB(PostgreSQL, MySQL), 메시지 브로커(Kafka), 분산 캐시(Redis Cluster)
PodDisruptionBudget선택 사항운영 환경에서 필수 권장

StatefulSet 완전 가이드

여기서는 StatefulSet 완전 가이드을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

PostgreSQL을 StatefulSet으로 배포하는 예제로 Headless Service, volumeClaimTemplates, 안정적 네트워크 ID, 운영 운용 패턴까지 단계별로 정리합니다.
postgres-statefulset.yamlYAML
# 1. Headless Service — Pod별 안정적 DNS (postgres-0.postgres-svc, postgres-1.postgres-svc)
apiVersion: v1
kind: Service
metadata:
  name: postgres-svc
  labels:
    app: postgres
spec:
  clusterIP: None          # Headless: Pod IP를 직접 DNS에 등록
  selector:
    app: postgres
  ports:
  - port: 5432
    name: postgres
---
# 2. StatefulSet
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: postgres
spec:
  serviceName: postgres-svc   # Headless Service 이름과 일치
  replicas: 3
  selector:
    matchLabels:
      app: postgres
  template:
    metadata:
      labels:
        app: postgres
    spec:
      containers:
      - name: postgres
        image: postgres:16-alpine
        ports:
        - containerPort: 5432
        env:
        - name: POSTGRES_PASSWORD
          valueFrom:
            secretKeyRef:
              name: postgres-secret
              key: password
        - name: PGDATA
          value: /var/lib/postgresql/data/pgdata
        resources:
          requests: { cpu: 250m, memory: 512Mi }
          limits:   { cpu: 1,    memory: 2Gi   }
        volumeMounts:
        - name: postgres-data
          mountPath: /var/lib/postgresql/data
        readinessProbe:
          exec:
            command: ["pg_isready", "-U", "postgres"]
          initialDelaySeconds: 5
          periodSeconds: 10
        livenessProbe:
          exec:
            command: ["pg_isready", "-U", "postgres"]
          initialDelaySeconds: 30
          periodSeconds: 20
  # 3. volumeClaimTemplates — Pod마다 독립 PVC 자동 생성
  volumeClaimTemplates:
  - metadata:
      name: postgres-data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: standard
      resources:
        requests:
          storage: 10Gi
ops-commands.shBASH
# 배포
kubectl apply -f postgres-statefulset.yaml

# Pod 상태 확인 (순서 보장: postgres-0 먼저 Running 후 postgres-1 생성)
kubectl get pods -l app=postgres -w

# Pod별 고정 DNS 확인
# postgres-0.postgres-svc.default.svc.cluster.local
# postgres-1.postgres-svc.default.svc.cluster.local
kubectl exec -it postgres-0 -- psql -U postgres -c "SELECT inet_server_addr();"

# PVC는 Pod 삭제 후에도 유지됨
kubectl delete pod postgres-2
kubectl get pvc | grep postgres      # postgres-data-postgres-2 그대로 존재

# 스케일 다운 (2→1→0 역순 삭제)
kubectl scale statefulset postgres --replicas=1

# 롤링 업데이트 (updateStrategy: RollingUpdate, 역순 2→1→0)
kubectl set image statefulset/postgres postgres=postgres:17-alpine
kubectl rollout status statefulset/postgres

# 특정 Pod만 재시작 (PVC 유지)
kubectl delete pod postgres-1
pdb.yamlYAML
# PodDisruptionBudget — 노드 드레인 등 자발적 중단 시 최소 2개 유지
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: postgres-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: postgres
kafka-statefulset.yamlYAML
# Kafka StatefulSet 예시 — Broker별 고유 ID, 독립 스토리지
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: kafka
spec:
  serviceName: kafka-svc
  replicas: 3
  selector:
    matchLabels:
      app: kafka
  template:
    metadata:
      labels:
        app: kafka
    spec:
      containers:
      - name: kafka
        image: bitnami/kafka:3.7
        env:
        - name: KAFKA_BROKER_ID            # Pod 순번으로 Broker ID 결정
          valueFrom:
            fieldRef:
              fieldPath: metadata.annotations['statefulset.kubernetes.io/pod-name']
        - name: KAFKA_ADVERTISED_LISTENERS
          value: "PLAINTEXT://$(POD_NAME).kafka-svc:9092"
        - name: POD_NAME
          valueFrom:
            fieldRef:
              fieldPath: metadata.name
        volumeMounts:
        - name: kafka-data
          mountPath: /bitnami/kafka
  volumeClaimTemplates:
  - metadata:
      name: kafka-data
    spec:
      accessModes: ["ReadWriteOnce"]
      resources:
        requests:
          storage: 20Gi

StatefulSet 배포 워크플로

여기서는 StatefulSet 배포 워크플로을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

StatefulSet을 프로덕션에 안전하게 배포하는 6단계 절차입니다. Namespace → StorageClass → Secret → Headless Service → StatefulSet → PodDisruptionBudget 순서로 진행합니다.
01-namespace.shBASH
# 1. Namespace 및 ResourceQuota 설정
kubectl create namespace database

kubectl apply -f - <<EOF
apiVersion: v1
kind: ResourceQuota
metadata:
  name: db-quota
  namespace: database
spec:
  hard:
    requests.cpu: "4"
    requests.memory: 8Gi
    limits.cpu: "8"
    limits.memory: 16Gi
    persistentvolumeclaims: "10"
EOF
02-storageclass.yamlYAML
# 2. StorageClass — SSD 기반 동적 프로비저닝
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast-ssd
provisioner: kubernetes.io/gce-pd   # GKE: gce-pd | EKS: ebs.csi.aws.com | NCP: nks-block-storage
parameters:
  type: pd-ssd
  replication-type: regional-pd
reclaimPolicy: Retain               # Pod/PVC 삭제 후에도 디스크 유지
allowVolumeExpansion: true          # 온라인 볼륨 확장 허용
volumeBindingMode: WaitForFirstConsumer
03-secret.shBASH
# 3. Secret 생성 (base64 인코딩 자동 처리)
kubectl create secret generic postgres-secret \
  --from-literal=POSTGRES_PASSWORD="$(openssl rand -base64 24)" \
  --from-literal=REPLICATION_PASSWORD="$(openssl rand -base64 24)" \
  --namespace=database

# 확인
kubectl get secret postgres-secret -n database -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
04-deploy.shBASH
# 4. Headless Service → StatefulSet 순서로 배포 (순서 필수)
kubectl apply -f postgres-headless-svc.yaml -n database
kubectl apply -f postgres-statefulset.yaml   -n database

# 5. Pod 기동 순서 모니터링 (postgres-0 Running 후 postgres-1 시작)
kubectl get pods -n database -l app=postgres -w

# Ready 상태 확인
kubectl rollout status statefulset/postgres -n database

# 6. PodDisruptionBudget 적용
kubectl apply -f postgres-pdb.yaml -n database

# 전체 리소스 확인
kubectl get all,pvc,pdb -n database -l app=postgres
05-scale-update.shBASH
# ── 스케일 업/다운 ────────────────────────────────────────────
# 스케일 업: 3 → 5 (3, 4 순서로 순차 생성)
kubectl scale statefulset postgres --replicas=5 -n database
kubectl get pods -n database -l app=postgres -w

# 스케일 다운: 5 → 3 (4, 3 순서로 역순 삭제, PVC는 유지)
kubectl scale statefulset postgres --replicas=3 -n database

# PVC는 삭제되지 않음 — 수동 정리 필요 시
kubectl delete pvc postgres-data-postgres-3 -n database
kubectl delete pvc postgres-data-postgres-4 -n database

# ── 롤링 업데이트 ─────────────────────────────────────────────
# 이미지 업데이트 (역순: pod-2 → pod-1 → pod-0)
kubectl set image statefulset/postgres \
  postgres=postgres:17-alpine -n database

# 업데이트 진행 상황
kubectl rollout status statefulset/postgres -n database

# 업데이트 일시 중단 (카나리 배포 패턴)
kubectl patch statefulset postgres -n database \
  -p '{"spec":{"updateStrategy":{"rollingUpdate":{"partition":2}}}}'
# partition=2 → pod-2만 새 버전, pod-0·pod-1은 유지
# 검증 후 partition=0으로 전체 업데이트

# ── 롤백 ─────────────────────────────────────────────────────
kubectl rollout undo statefulset/postgres -n database
kubectl rollout history statefulset/postgres -n database
06-disaster-recovery.shBASH
# ── 장애 복구 시나리오 ───────────────────────────────────────
# Pod 강제 재시작 (PVC 유지, 데이터 보존)
kubectl delete pod postgres-1 -n database
kubectl get pods -n database -w    # 자동 재생성 확인

# Pod가 Pending 상태일 때 진단
kubectl describe pod postgres-0 -n database
kubectl get events -n database --sort-by='.lastTimestamp'

# PVC 볼륨 상태 확인
kubectl get pvc -n database
kubectl describe pvc postgres-data-postgres-0 -n database

# 노드 장애 시 강제 Pod 재스케줄 (taint 제거가 안 될 때)
kubectl delete pod postgres-0 -n database --grace-period=0 --force

# 특정 Pod 로그 확인 (이전 컨테이너 로그 포함)
kubectl logs postgres-0 -n database --previous
kubectl logs postgres-0 -n database -f

실전 샘플

여기서는 실전 샘플을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Redis Sentinel, MySQL Primary-Replica, Elasticsearch 세 가지 프로덕션 패턴을 정리합니다. 각 샘플은 Headless Service + StatefulSet + volumeClaimTemplates 구조를 기반으로 합니다.
redis-sentinel.yamlYAML
# Redis Sentinel — HA 구성 (Primary 1 + Replica 2)
# Headless Service
apiVersion: v1
kind: Service
metadata:
  name: redis-svc
  namespace: cache
spec:
  clusterIP: None
  selector:
    app: redis
  ports:
  - port: 6379
    name: redis
  - port: 26379
    name: sentinel
---
# Client용 ClusterIP Service (Primary만 연결)
apiVersion: v1
kind: Service
metadata:
  name: redis-primary
  namespace: cache
spec:
  selector:
    app: redis
    role: primary
  ports:
  - port: 6379
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: redis
  namespace: cache
spec:
  serviceName: redis-svc
  replicas: 3
  selector:
    matchLabels:
      app: redis
  template:
    metadata:
      labels:
        app: redis
    spec:
      initContainers:
      - name: init-redis
        image: redis:7-alpine
        command:
        - bash
        - -c
        - |
          # pod-0이면 Primary, 나머지는 Replica 설정
          [[ $(hostname) == *-0 ]] && ROLE=primary || ROLE=replica
          if [[ $ROLE == primary ]]; then
            cp /mnt/config/redis-primary.conf /etc/redis/redis.conf
          else
            # primary 호스트명 주입
            sed "s/PRIMARY_HOST/redis-0.redis-svc/" \
              /mnt/config/redis-replica.conf > /etc/redis/redis.conf
          fi
        volumeMounts:
        - name: config-template
          mountPath: /mnt/config
        - name: redis-config
          mountPath: /etc/redis
      containers:
      - name: redis
        image: redis:7-alpine
        command: ["redis-server", "/etc/redis/redis.conf"]
        ports:
        - containerPort: 6379
        resources:
          requests: { cpu: 100m, memory: 256Mi }
          limits:   { cpu: 500m, memory: 1Gi }
        volumeMounts:
        - name: redis-data
          mountPath: /data
        - name: redis-config
          mountPath: /etc/redis
        readinessProbe:
          exec:
            command: ["redis-cli", "ping"]
          initialDelaySeconds: 5
          periodSeconds: 5
      - name: sentinel
        image: redis:7-alpine
        command: ["redis-sentinel", "/etc/sentinel/sentinel.conf"]
        ports:
        - containerPort: 26379
        volumeMounts:
        - name: sentinel-config
          mountPath: /etc/sentinel
      volumes:
      - name: config-template
        configMap:
          name: redis-config
      - name: redis-config
        emptyDir: {}
      - name: sentinel-config
        configMap:
          name: sentinel-config
  volumeClaimTemplates:
  - metadata:
      name: redis-data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources:
        requests:
          storage: 5Gi
mysql-primary-replica.yamlYAML
# MySQL Primary-Replica — initContainer로 역할 자동 분기
apiVersion: v1
kind: Service
metadata:
  name: mysql-svc
  namespace: database
spec:
  clusterIP: None
  selector:
    app: mysql
  ports:
  - port: 3306
---
apiVersion: v1
kind: Service
metadata:
  name: mysql-read
  namespace: database
spec:
  selector:
    app: mysql
  ports:
  - port: 3306
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: mysql
  namespace: database
spec:
  serviceName: mysql-svc
  replicas: 3
  selector:
    matchLabels:
      app: mysql
  template:
    metadata:
      labels:
        app: mysql
    spec:
      initContainers:
      - name: init-mysql
        image: mysql:8.4
        command:
        - bash
        - -c
        - |
          # Pod 순번으로 server-id 결정, pod-0은 Primary
          ORDINAL=$(hostname | awk -F- '{print $NF}')
          echo [mysqld] > /mnt/conf.d/server-id.cnf
          echo server-id=$((100 + $ORDINAL)) >> /mnt/conf.d/server-id.cnf
          if [[ $ORDINAL -eq 0 ]]; then
            cp /mnt/config/primary.cnf /mnt/conf.d/
          else
            cp /mnt/config/replica.cnf /mnt/conf.d/
          fi
        volumeMounts:
        - name: conf
          mountPath: /mnt/conf.d
        - name: config-map
          mountPath: /mnt/config
      - name: clone-mysql
        image: gcr.io/google-samples/xtrabackup:1.0
        command:
        - bash
        - -c
        - |
          # pod-0이 아닌 경우 이전 Pod에서 데이터 복제
          ORDINAL=$(hostname | awk -F- '{print $NF}')
          [[ $ORDINAL -eq 0 ]] && exit 0
          PEER=mysql-$(($ORDINAL - 1)).mysql-svc
          ncat --recv-only $PEER 3307 | xbstream -x -C /var/lib/mysql
          xtrabackup --prepare --target-dir=/var/lib/mysql
        volumeMounts:
        - name: data
          mountPath: /var/lib/mysql
          subPath: mysql
        - name: conf
          mountPath: /etc/mysql/conf.d
      containers:
      - name: mysql
        image: mysql:8.4
        env:
        - name: MYSQL_ROOT_PASSWORD
          valueFrom:
            secretKeyRef:
              name: mysql-secret
              key: ROOT_PASSWORD
        ports:
        - containerPort: 3306
        resources:
          requests: { cpu: 500m, memory: 1Gi }
          limits:   { cpu: 2,    memory: 4Gi }
        volumeMounts:
        - name: data
          mountPath: /var/lib/mysql
          subPath: mysql
        - name: conf
          mountPath: /etc/mysql/conf.d
        readinessProbe:
          exec:
            command: ["mysqladmin", "ping", "-uroot", "-p$(MYSQL_ROOT_PASSWORD)"]
          initialDelaySeconds: 30
          periodSeconds: 10
      - name: xtrabackup
        image: gcr.io/google-samples/xtrabackup:1.0
        ports:
        - containerPort: 3307
        command:
        - bash
        - -c
        - |
          # Replica 클론 요청 대기 (ncat 서버)
          cd /var/lib/mysql
          if [[ -f xtrabackup_slave_info && "x$(<xtrabackup_slave_info)" != "x" ]]; then
            cat xtrabackup_slave_info | sed -E 's/;*s*$//; s/MASTER/SOURCE/' \
              | mysql -uroot -p${MYSQL_ROOT_PASSWORD}
            rm -f xtrabackup_slave_info xtrabackup_binlog_info
          fi
          exec ncat --listen --keep-open --send-only --max-conns=1 3307 -c \
            "xtrabackup --backup --slave-info --stream=xbstream --host=127.0.0.1 --user=root --password=${MYSQL_ROOT_PASSWORD}"
        volumeMounts:
        - name: data
          mountPath: /var/lib/mysql
          subPath: mysql
        - name: conf
          mountPath: /etc/mysql/conf.d
      volumes:
      - name: conf
        emptyDir: {}
      - name: config-map
        configMap:
          name: mysql-config
  volumeClaimTemplates:
  - metadata:
      name: data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources:
        requests:
          storage: 20Gi
elasticsearch.yamlYAML
# Elasticsearch StatefulSet — 3노드 클러스터
apiVersion: v1
kind: Service
metadata:
  name: es-svc
  namespace: search
spec:
  clusterIP: None
  selector:
    app: elasticsearch
  ports:
  - port: 9200
    name: http
  - port: 9300
    name: transport
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: elasticsearch
  namespace: search
spec:
  serviceName: es-svc
  replicas: 3
  selector:
    matchLabels:
      app: elasticsearch
  template:
    metadata:
      labels:
        app: elasticsearch
    spec:
      initContainers:
      - name: fix-permissions
        image: busybox
        command: ["sh", "-c", "chown -R 1000:1000 /usr/share/elasticsearch/data"]
        securityContext:
          privileged: true
        volumeMounts:
        - name: es-data
          mountPath: /usr/share/elasticsearch/data
      - name: increase-vm-max-map
        image: busybox
        command: ["sysctl", "-w", "vm.max_map_count=262144"]
        securityContext:
          privileged: true
      containers:
      - name: elasticsearch
        image: docker.elastic.co/elasticsearch/elasticsearch:8.13.0
        env:
        - name: cluster.name
          value: "k8s-es-cluster"
        - name: node.name
          valueFrom:
            fieldRef:
              fieldPath: metadata.name
        - name: discovery.seed_hosts
          value: "elasticsearch-0.es-svc,elasticsearch-1.es-svc,elasticsearch-2.es-svc"
        - name: cluster.initial_master_nodes
          value: "elasticsearch-0,elasticsearch-1,elasticsearch-2"
        - name: ES_JAVA_OPTS
          value: "-Xms2g -Xmx2g"
        - name: xpack.security.enabled
          value: "true"
        - name: ELASTIC_PASSWORD
          valueFrom:
            secretKeyRef:
              name: es-secret
              key: password
        ports:
        - containerPort: 9200
          name: http
        - containerPort: 9300
          name: transport
        resources:
          requests: { cpu: 500m, memory: 3Gi }
          limits:   { cpu: 2,    memory: 5Gi }
        volumeMounts:
        - name: es-data
          mountPath: /usr/share/elasticsearch/data
        readinessProbe:
          httpGet:
            path: /_cluster/health?local=true
            port: 9200
            httpHeaders:
            - name: Authorization
              value: "Basic ZWxhc3RpYzpjaGFuZ2VtZQ=="
          initialDelaySeconds: 30
          periodSeconds: 10
  volumeClaimTemplates:
  - metadata:
      name: es-data
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources:
        requests:
          storage: 30Gi
verify-all.shBASH
# ── 공통 배포 검증 스크립트 ──────────────────────────────────

# 1. 모든 Pod Running 확인
kubectl get pods -A -l app in (postgres,redis,mysql,elasticsearch) \
  --field-selector=status.phase!=Running

# 2. PVC Bound 확인
kubectl get pvc -A | grep -v Bound

# 3. StatefulSet 롤아웃 완료 확인
for sts in postgres redis mysql elasticsearch; do
  echo "=== $sts ==="
  kubectl rollout status statefulset/$sts --timeout=120s 2>/dev/null || true
done

# 4. Endpoint 통신 확인
# PostgreSQL
kubectl exec -it postgres-0 -n database -- \
  psql -U postgres -c "SELECT version();"

# Redis
kubectl exec -it redis-0 -n cache -- redis-cli ping
kubectl exec -it redis-0 -n cache -- \
  redis-cli info replication | grep -E "role|connected_slaves"

# Elasticsearch
kubectl exec -it elasticsearch-0 -n search -- \
  curl -s -u elastic:$ELASTIC_PASSWORD \
  http://localhost:9200/_cluster/health?pretty | jq '.status'

# 5. 리소스 사용량
kubectl top pods -A -l app in (postgres,redis,mysql,elasticsearch)

StatefulSet 심화 — Parallel · 볼륨 확장 · 스냅샷

여기서는 StatefulSet 심화 — Parallel · 볼륨 확장 · 스냅샷을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

기본 StatefulSet 운영에 익숙해졌다면 podManagementPolicy로 기동 순서를 제어하고, PVC를 무중단으로 확장하고, VolumeSnapshot으로 데이터를 백업·복구하는 실무 패턴을 알아야 합니다. 특히 volumeClaimTemplates는 StatefulSet 생성 후에는 직접 수정할 수 없다는 제약이 자주 실수를 유발합니다.
statefulset-parallel.yamlYAML
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: elasticsearch
spec:
  serviceName: elasticsearch-svc
  replicas: 5
  podManagementPolicy: Parallel   # 5개 Pod을 동시에 기동 — 순서 대기 없이 스케일 속도 향상
  selector:
    matchLabels: { app: elasticsearch }
  template:
    metadata:
      labels: { app: elasticsearch }
    spec:
      containers:
      - name: elasticsearch
        image: elasticsearch:8.13.0
        env:
        - name: discovery.seed_hosts
          value: "elasticsearch-0.elasticsearch-svc,elasticsearch-1.elasticsearch-svc"
        - name: cluster.initial_master_nodes
          value: "elasticsearch-0,elasticsearch-1,elasticsearch-2"
        volumeMounts:
        - name: es-data
          mountPath: /usr/share/elasticsearch/data
  volumeClaimTemplates:
  - metadata: { name: es-data }
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources: { requests: { storage: 100Gi } }
pvc-online-expand.shBASH
# ── PVC 온라인 확장 (StorageClass에 allowVolumeExpansion: true 필수) ──
# ⚠️ volumeClaimTemplates는 StatefulSet 생성 후 직접 patch/edit 불가 (immutable field)
# → 이미 생성된 PVC들을 개별적으로 patch 해야 함

# 1. 각 PVC 용량을 직접 확장 (10Gi → 50Gi)
for i in 0 1 2; do
  kubectl patch pvc postgres-data-postgres-$i -n database \
    -p '{"spec":{"resources":{"requests":{"storage":"50Gi"}}}}'
done

# 2. 확장 진행 상태 확인 (FileSystemResizePending → 완료 시 사라짐)
kubectl get pvc -n database -w
kubectl describe pvc postgres-data-postgres-0 -n database | grep -A3 Conditions

# 3. StatefulSet manifest의 volumeClaimTemplates 용량 값도 함께 맞춰 반영
#    (다음 재생성/신규 replica 추가 시 새 PVC가 같은 크기로 생성되도록)
#    spec.volumeClaimTemplates 필드 자체는 수정 불가하므로,
#    변경이 필요하면 --cascade=orphan으로 Pod을 보존한 채 StatefulSet만 재생성한다
kubectl delete statefulset postgres -n database --cascade=orphan
kubectl apply -f postgres-statefulset.yaml -n database   # 용량 값을 50Gi로 수정 후 재적용
volumesnapshot-backup.yamlYAML
# ── VolumeSnapshot 기반 백업 & 복구 (CSI 드라이버 + VolumeSnapshotClass 필요) ──
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
  name: csi-snapclass
driver: ebs.csi.aws.com          # GKE: pd.csi.storage.gke.io | NCP: nks-block-storage
deletionPolicy: Retain            # 원본 PVC 삭제되어도 스냅샷 유지
---
# 1. 스냅샷 생성 — 배포 전 / 대규모 마이그레이션 전 백업
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: postgres-snap-before-upgrade
  namespace: database
spec:
  volumeSnapshotClassName: csi-snapclass
  source:
    persistentVolumeClaimName: postgres-data-postgres-0
---
# 2. 스냅샷으로부터 새 PVC 복구 — 장애 발생 시 별도 이름으로 복원해 데이터 검증 후 교체
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: postgres-data-restored
  namespace: database
spec:
  storageClassName: fast-ssd
  dataSource:
    name: postgres-snap-before-upgrade
    kind: VolumeSnapshot
    apiGroup: snapshot.storage.k8s.io
  accessModes: ["ReadWriteOnce"]
  resources:
    requests:
      storage: 50Gi
snapshot-ops.shBASH
# 스냅샷 상태 확인 (readyToUse: true 될 때까지 대기)
kubectl get volumesnapshot -n database -w
kubectl describe volumesnapshot postgres-snap-before-upgrade -n database

# 복구된 PVC를 검증용 Pod에 마운트해 데이터 확인 후, 문제없으면 기존 PVC와 교체
kubectl run pg-verify --image=postgres:16-alpine --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"pg-verify","image":"postgres:16-alpine","command":["sleep","3600"],"volumeMounts":[{"name":"data","mountPath":"/data"}]}],"volumes":[{"name":"data","persistentVolumeClaim":{"claimName":"postgres-data-restored"}}]}}' \
  -n database
podManagementPolicy동작적합한 워크로드
OrderedReady (기본값)Pod-0이 Running & Ready가 되어야 Pod-1을 생성. 삭제도 역순으로 하나씩PostgreSQL Primary-Replica처럼 기동 순서가 데이터 정합성에 영향을 주는 경우
Parallel모든 Pod을 동시에 생성/삭제. 순서 보장 없음, 대신 훨씬 빠름Elasticsearch·Cassandra·Kafka처럼 각 노드가 독립적으로 클러스터에 join하는 경우

DaemonSet 완전 가이드

여기서는 DaemonSet 완전 가이드을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

DaemonSet은 클러스터의 모든(또는 라벨로 선택된) 노드에 Pod을 정확히 1개씩 배포합니다. 새 노드가 추가되면 자동으로 Pod이 생성되고, 노드가 제거되면 함께 정리됩니다. 로그 수집기, 모니터링 에이전트, CNI/CSI 플러그인처럼 "노드마다 하나씩 있어야 하는" 인프라 컴포넌트에 적합합니다.
다이어그램 렌더링 중…
fluent-bit-daemonset.yamlYAML
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: fluent-bit
  namespace: logging
spec:
  selector:
    matchLabels: { app: fluent-bit }
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1        # 한 번에 노드 1개씩만 재시작 — 로그 유실 구간 최소화
  template:
    metadata:
      labels: { app: fluent-bit }
    spec:
      tolerations:              # control-plane 노드에도 배포해 마스터 로그까지 수집
      - key: node-role.kubernetes.io/control-plane
        effect: NoSchedule
      serviceAccountName: fluent-bit
      containers:
      - name: fluent-bit
        image: fluent/fluent-bit:3.0
        resources:
          requests: { cpu: 100m, memory: 128Mi }   # 노드 수만큼 곱해지므로 최소로 유지
          limits:   { cpu: 300m, memory: 256Mi }
        volumeMounts:
        - name: varlog
          mountPath: /var/log            # 노드의 로그 디렉터리를 직접 마운트
          readOnly: true
        - name: dockercontainers
          mountPath: /var/lib/docker/containers
          readOnly: true
      volumes:
      - name: varlog
        hostPath: { path: /var/log }
      - name: dockercontainers
        hostPath: { path: /var/lib/docker/containers }
daemonset-ops.shBASH
# 노드별 배포 현황 확인 — READY가 노드 수와 일치해야 함
kubectl get daemonset fluent-bit -n logging
kubectl get pods -n logging -l app=fluent-bit -o wide

# 특정 노드에서만 실행되도록 제한 (예: GPU 노드에만 device-plugin 배포)
kubectl label node gpu-node-1 workload=gpu
# spec.template.spec.nodeSelector: { workload: gpu } 를 매니페스트에 추가

# 롤링 업데이트
kubectl set image daemonset/fluent-bit fluent-bit=fluent/fluent-bit:3.1 -n logging
kubectl rollout status daemonset/fluent-bit -n logging
kubectl rollout history daemonset/fluent-bit -n logging

# 신규 노드 추가 시 자동 스케줄 확인
kubectl get nodes -w
특성DaemonSetDeployment
replicas 지정불가 — 대상 노드 수만큼 자동 결정명시적으로 지정
스케줄링노드당 정확히 1개, nodeSelector/affinity로 대상 노드 제한 가능스케줄러가 리소스 여유가 있는 노드에 자유 배치
Control-plane 노드 배포tolerations로 NoSchedule taint를 허용해야 배포됨기본적으로 배제됨
업데이트 전략RollingUpdate(maxUnavailable) 또는 OnDeleteRollingUpdate(maxSurge/maxUnavailable) 또는 Recreate
대표 사례Fluent Bit, Node Exporter, Calico/Cilium, NVIDIA device plugin웹 서버, API 서버 등 일반 애플리케이션

CronJob & Job 완전 가이드

여기서는 CronJob & Job 완전 가이드을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Job은 완료가 보장되어야 하는 일회성 작업(배치, 마이그레이션)을, CronJob은 그 Job을 cron 스케줄에 따라 주기적으로 생성합니다. 실패 시 재시도 횟수(backoffLimit), 최대 실행 시간(activeDeadlineSeconds), 동시 실행 정책(concurrencyPolicy)을 명확히 정의하지 않으면 배치 작업이 무한 재시도하거나 중복 실행되는 사고로 이어집니다.
db-migration-job.yamlYAML
apiVersion: batch/v1
kind: Job
metadata:
  name: db-migration-v2
  namespace: production
spec:
  backoffLimit: 3                 # 3회 실패하면 포기 (무한 재시도 방지)
  activeDeadlineSeconds: 600       # 10분 넘으면 강제 종료
  ttlSecondsAfterFinished: 3600    # 완료 1시간 후 자동 정리
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: migrate
        image: myapp-migrator:2.1.0
        command: ["./migrate.sh", "--target=v2"]
        resources:
          requests: { cpu: 200m, memory: 256Mi }
          limits:   { cpu: 500m, memory: 512Mi }
        envFrom:
        - secretRef: { name: db-credentials }
nightly-backup-cronjob.yamlYAML
apiVersion: batch/v1
kind: CronJob
metadata:
  name: postgres-nightly-backup
  namespace: database
spec:
  schedule: "0 3 * * *"            # 매일 새벽 3시 (클러스터 kube-controller-manager 시간대 기준)
  timeZone: "Asia/Seoul"           # K8s 1.27+ — CronJob 자체에 타임존 지정 가능
  concurrencyPolicy: Forbid        # 이전 백업이 아직 실행 중이면 이번 회차는 건너뜀
  startingDeadlineSeconds: 300     # 5분 안에 시작 못하면 이번 회차는 포기
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 5        # 실패 이력은 조금 더 오래 남겨 원인 분석
  jobTemplate:
    spec:
      backoffLimit: 2
      activeDeadlineSeconds: 1800
      template:
        spec:
          restartPolicy: OnFailure
          containers:
          - name: backup
            image: postgres:16-alpine
            command:
            - sh
            - -c
            - |
              pg_dump -h postgres-svc -U postgres mydb | \
              gzip > /backup/mydb-$(date +%Y%m%d).sql.gz
            volumeMounts:
            - name: backup-storage
              mountPath: /backup
          volumes:
          - name: backup-storage
            persistentVolumeClaim: { claimName: backup-pvc }
job-cronjob-ops.shBASH
# CronJob으로부터 즉시 1회 수동 실행 (스케줄 기다리지 않고 테스트)
kubectl create job --from=cronjob/postgres-nightly-backup manual-backup-test -n database

# 실행 이력 확인
kubectl get cronjob postgres-nightly-backup -n database
kubectl get jobs -n database -l job-name  --sort-by=.metadata.creationTimestamp

# 특정 Job의 Pod 로그 확인
kubectl logs -n database -l job-name=postgres-nightly-backup-28912345 --tail=100

# CronJob 일시 중지/재개 (배포 중이거나 장애 조사 중일 때)
kubectl patch cronjob postgres-nightly-backup -n database -p '{"spec":{"suspend":true}}'
kubectl patch cronjob postgres-nightly-backup -n database -p '{"spec":{"suspend":false}}'

# 실패한 Job 정리 (ttlSecondsAfterFinished 미설정 시 수동 정리)
kubectl delete job --field-selector status.successful=1 -n database
필드적용 대상역할
backoffLimitJob실패 시 재시도 최대 횟수 (기본 6) — 초과하면 Job이 Failed로 종료
activeDeadlineSecondsJob전체 실행 제한 시간 — 초과 시 강제 종료, 무한 실행 방지
restartPolicyJob PodNever(재시도 시 새 Pod 생성) 또는 OnFailure(같은 Pod 재시작) — Always 불가
ttlSecondsAfterFinishedJob완료 후 N초 뒤 Job/Pod 자동 삭제 — 완료된 Job이 계속 쌓이는 것 방지
concurrencyPolicyCronJobAllow(기본, 중복 허용) / Forbid(이전 실행 중이면 스킵) / Replace(이전 실행 취소 후 새로 시작)
startingDeadlineSecondsCronJob스케줄 시각을 놓쳤을 때(컨트롤러 다운 등) 몇 초까지 지연 실행을 허용할지
successfulJobsHistoryLimit / failedJobsHistoryLimitCronJob보관할 성공/실패 Job 이력 개수 — 디버깅용으로 최근 몇 개만 남김

Java Agent StatefulSet 배포

여기서는 Java Agent StatefulSet 배포을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Spring Boot 기반 Java AI Agent를 StatefulSet으로 배포하는 실전 가이드입니다. Agent 상태·벡터 인덱스를 저장하는 PVC와 감사 로그를 분리 보관하는 PVC 2개를 각 Pod 전용으로 프로비저닝합니다.
DockerfileDOCKERFILE
# ── Multi-stage build — Spring Boot AI Agent ──────────────────
FROM eclipse-temurin:21-jdk-alpine AS builder
WORKDIR /app
COPY gradlew build.gradle.kts settings.gradle.kts ./
COPY gradle ./gradle
RUN ./gradlew dependencies --no-daemon -q

COPY src ./src
RUN ./gradlew bootJar --no-daemon -q

# ── Runtime image ─────────────────────────────────────────────
FROM eclipse-temurin:21-jre-alpine
WORKDIR /app

# 데이터 · 로그 디렉터리 생성 (PVC 마운트 포인트)
RUN mkdir -p /app/data /app/logs && \
    addgroup -S agent && adduser -S agent -G agent && \
    chown -R agent:agent /app

COPY --from=builder /app/build/libs/*.jar agent.jar
USER agent

# PVC 볼륨 선언 (K8s가 마운트)
VOLUME ["/app/data", "/app/logs"]

EXPOSE 8080
ENTRYPOINT ["java", \
  "-XX:MaxRAMPercentage=75.0", \
  "-XX:+UseG1GC", \
  "-Dspring.profiles.active=prod", \
  "-Dlogging.file.path=/app/logs", \
  "-jar", "agent.jar"]
configmap.yamlYAML
apiVersion: v1
kind: ConfigMap
metadata:
  name: java-agent-config
  namespace: ai-agent
data:
  application-prod.yml: |
    spring:
      datasource:
        url: jdbc:h2:file:/app/data/agentdb;AUTO_SERVER=TRUE
        driver-class-name: org.h2.Driver
        username: agent
      jpa:
        hibernate:
          ddl-auto: update
        show-sql: false
      ai:
        openai:
          base-url: http://llm-gateway-svc:8080   # 내부 LLM Gateway
          api-key: ${OPENAI_API_KEY}

    agent:
      vector-store:
        path: /app/data/faiss-index          # FAISS 로컬 인덱스
        dimension: 1536
      memory:
        persist-path: /app/data/memory.db   # Agent 메모리 저장
        max-tokens: 4096

    logging:
      file:
        name: /app/logs/agent.log
        max-size: 100MB
        max-history: 90                       # 90일 보관
      pattern:
        file: "%d{yyyy-MM-dd HH:mm:ss} [%X{traceId}] %-5level %logger{36} - %msg%n"

    management:
      endpoints:
        web:
          exposure:
            include: health,metrics,prometheus
      metrics:
        export:
          prometheus:
            enabled: true
secret.shBASH
# Secret 생성 (API 키, DB 패스워드)
kubectl create namespace ai-agent

kubectl create secret generic java-agent-secret \
  --from-literal=OPENAI_API_KEY="sk-xxxx" \
  --from-literal=SPRING_DATASOURCE_PASSWORD="$(openssl rand -base64 16)" \
  --namespace=ai-agent

# 컨테이너 레지스트리 인증 (Private registry)
kubectl create secret docker-registry regcred \
  --docker-server=registry.aidevops.kr \
  --docker-username=deploy \
  --docker-password="<token>" \
  --namespace=ai-agent
statefulset.yamlYAML
# ── Headless Service (Pod 간 DNS) ────────────────────────────
apiVersion: v1
kind: Service
metadata:
  name: java-agent-svc
  namespace: ai-agent
spec:
  clusterIP: None
  selector:
    app: java-agent
  ports:
  - port: 8080
    name: http
---
# ── ClusterIP Service (내부 API 호출용) ──────────────────────
apiVersion: v1
kind: Service
metadata:
  name: java-agent-api
  namespace: ai-agent
spec:
  selector:
    app: java-agent
  ports:
  - port: 80
    targetPort: 8080
    name: http
---
# ── StatefulSet ───────────────────────────────────────────────
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: java-agent
  namespace: ai-agent
spec:
  serviceName: java-agent-svc       # Headless Service 이름
  replicas: 2
  selector:
    matchLabels:
      app: java-agent
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      partition: 0                  # 전체 롤링 업데이트
  template:
    metadata:
      labels:
        app: java-agent
    spec:
      terminationGracePeriodSeconds: 60
      imagePullSecrets:
      - name: regcred
      initContainers:
      # data 디렉터리 권한 초기화 (PVC 마운트 시 root 소유 문제 방지)
      - name: init-permissions
        image: busybox:1.36
        command: ["sh", "-c", "chown -R 1000:1000 /app/data /app/logs"]
        volumeMounts:
        - name: agent-state
          mountPath: /app/data
        - name: agent-logs
          mountPath: /app/logs
      containers:
      - name: java-agent
        image: registry.aidevops.kr/java-agent:1.0.0
        ports:
        - containerPort: 8080
          name: http
        env:
        - name: POD_NAME
          valueFrom:
            fieldRef:
              fieldPath: metadata.name
        - name: POD_NAMESPACE
          valueFrom:
            fieldRef:
              fieldPath: metadata.namespace
        - name: OPENAI_API_KEY
          valueFrom:
            secretKeyRef:
              name: java-agent-secret
              key: OPENAI_API_KEY
        - name: SPRING_DATASOURCE_PASSWORD
          valueFrom:
            secretKeyRef:
              name: java-agent-secret
              key: SPRING_DATASOURCE_PASSWORD
        # Pod 고유 ID를 Agent ID로 활용 (java-agent-0, java-agent-1)
        - name: AGENT_ID
          value: "$(POD_NAME)"
        envFrom:
        - configMapRef:
            name: java-agent-config   # application-prod.yml 마운트용 별도 ConfigMap
        resources:
          requests:
            cpu: 500m
            memory: 1Gi
          limits:
            cpu: 2
            memory: 3Gi
        volumeMounts:
        - name: agent-state           # PVC 1 — 에이전트 상태 · 벡터 인덱스
          mountPath: /app/data
        - name: agent-logs            # PVC 2 — 감사/운영 로그
          mountPath: /app/logs
        - name: app-config
          mountPath: /app/config
          readOnly: true
        readinessProbe:
          httpGet:
            path: /actuator/health/readiness
            port: 8080
          initialDelaySeconds: 30
          periodSeconds: 10
          failureThreshold: 3
        livenessProbe:
          httpGet:
            path: /actuator/health/liveness
            port: 8080
          initialDelaySeconds: 60
          periodSeconds: 20
          failureThreshold: 4
        lifecycle:
          preStop:
            exec:
              # Graceful shutdown — 진행 중인 Agent 작업 완료 대기
              command: ["sh", "-c", "sleep 10"]
      volumes:
      - name: app-config
        configMap:
          name: java-agent-config
  # ── PVC 2개 — volumeClaimTemplates ────────────────────────
  volumeClaimTemplates:
  - metadata:
      name: agent-state
      labels:
        app: java-agent
        volume: state
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: fast-ssd
      resources:
        requests:
          storage: 20Gi
  - metadata:
      name: agent-logs
      labels:
        app: java-agent
        volume: logs
    spec:
      accessModes: ["ReadWriteOnce"]
      storageClassName: nfs-client     # 로그는 NFS로 중앙화 가능
      resources:
        requests:
          storage: 10Gi
ingress.yamlYAML
# Ingress — 외부에서 Java Agent API 접근
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: java-agent-ingress
  namespace: ai-agent
  annotations:
    nginx.ingress.kubernetes.io/rewrite-target: /
    nginx.ingress.kubernetes.io/ssl-redirect: "true"
    nginx.ingress.kubernetes.io/proxy-read-timeout: "300"    # Agent 응답 대기
    nginx.ingress.kubernetes.io/proxy-send-timeout: "300"
    cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
  ingressClassName: nginx
  tls:
  - hosts:
    - agent-api.aidevops.kr
    secretName: agent-api-tls
  rules:
  - host: agent-api.aidevops.kr
    http:
      paths:
      - path: /api/agent
        pathType: Prefix
        backend:
          service:
            name: java-agent-api
            port:
              number: 80
pdb.yamlYAML
# PodDisruptionBudget — 최소 1개 Pod 항상 유지
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: java-agent-pdb
  namespace: ai-agent
spec:
  minAvailable: 1
  selector:
    matchLabels:
      app: java-agent
deploy-and-verify.shBASH
# ── 전체 배포 순서 ────────────────────────────────────────────
kubectl apply -f configmap.yaml
kubectl apply -f statefulset.yaml    # Headless SVC + ClusterIP SVC + StatefulSet
kubectl apply -f ingress.yaml
kubectl apply -f pdb.yaml

# Pod 기동 확인 (java-agent-0 먼저 Running 후 java-agent-1)
kubectl get pods -n ai-agent -l app=java-agent -w

# PVC 2개씩 생성 확인 (Pod당 agent-state + agent-logs)
kubectl get pvc -n ai-agent
# 예상 출력:
# agent-state-java-agent-0   Bound   20Gi
# agent-logs-java-agent-0    Bound   10Gi
# agent-state-java-agent-1   Bound   20Gi
# agent-logs-java-agent-1    Bound   10Gi

# Actuator 헬스 확인
kubectl exec -it java-agent-0 -n ai-agent -- \
  curl -s http://localhost:8080/actuator/health | jq .

# 데이터 디렉터리 확인 (FAISS 인덱스, H2 DB)
kubectl exec -it java-agent-0 -n ai-agent -- ls -lh /app/data/
kubectl exec -it java-agent-0 -n ai-agent -- ls -lh /app/logs/

# 로그 실시간 확인
kubectl logs java-agent-0 -n ai-agent -f

# ── 롤링 업데이트 ─────────────────────────────────────────────
kubectl set image statefulset/java-agent \
  java-agent=registry.aidevops.kr/java-agent:1.1.0 \
  -n ai-agent
kubectl rollout status statefulset/java-agent -n ai-agent

# ── Pod 재시작 (PVC 데이터 보존) ──────────────────────────────
kubectl delete pod java-agent-1 -n ai-agent
kubectl get pods -n ai-agent -w   # 자동 재생성 확인

# ── PVC 용량 확인 ─────────────────────────────────────────────
kubectl exec -it java-agent-0 -n ai-agent -- df -h /app/data /app/logs
PVC 이름마운트 경로용도크기
agent-state/app/dataH2 embedded DB, 벡터 인덱스(FAISS), 에이전트 메모리, 파일 업로드20 Gi
agent-logs/app/logs감사 로그, 트레이스 로그 (Logback 아카이브, 90일 보관)10 Gi

GPU 워크로드 & LLM 추론 서버 배포

여기서는 GPU 워크로드 & LLM 추론 서버 배포을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

AI Agent/LLM 서빙 워크로드는 GPU 리소스를 정확히 요청하고, 일반 CPU Pod와 같은 노드에 스케줄되지 않도록 격리해야 합니다. NVIDIA device plugin이 GPU를 노드 리소스(nvidia.com/gpu)로 노출하면, requests/limits에 정수 단위로만 지정할 수 있습니다 — GPU는 CPU/메모리처럼 분할(fractional) 요청이 불가능합니다.
BASH
# NVIDIA device plugin 설치 — GPU 노드를 nvidia.com/gpu 리소스로 노출
kubectl create -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/main/deployments/static/nvidia-device-plugin.yml

# GPU 노드 확인
kubectl get nodes -o json | jq '.items[].status.capacity."nvidia.com/gpu"'

# GPU 노드에 전용 taint 부여 — 일반 Pod가 실수로 스케줄되는 것을 방지
kubectl taint nodes gpu-node-1 nvidia.com/gpu=present:NoSchedule
llm-inference-deploy.yamlYAML
apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-inference
  namespace: ai-agent
spec:
  replicas: 2
  selector:
    matchLabels: { app: llm-inference }
  template:
    metadata:
      labels: { app: llm-inference }
    spec:
      nodeSelector:
        gpu-type: a10g                 # GPU 종류별로 노드 풀을 분리해 라벨링
      tolerations:
      - key: nvidia.com/gpu
        operator: Equal
        value: present
        effect: NoSchedule
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args: ["--model", "/models/llama-3-8b", "--gpu-memory-utilization", "0.9"]
        resources:
          requests:
            cpu: "4"
            memory: 24Gi
            nvidia.com/gpu: "1"        # GPU는 정수 단위만 요청 가능 — 분할 불가
          limits:
            cpu: "8"
            memory: 32Gi
            nvidia.com/gpu: "1"
        volumeMounts:
        - name: model-cache
          mountPath: /models
        readinessProbe:                 # 모델 로딩(수십 초~수 분) 완료 전 트래픽 차단
          httpGet: { path: /health, port: 8000 }
          initialDelaySeconds: 60
          periodSeconds: 10
          failureThreshold: 30
        livenessProbe:
          httpGet: { path: /health, port: 8000 }
          initialDelaySeconds: 120
          periodSeconds: 30
      volumes:
      - name: model-cache
        persistentVolumeClaim:
          claimName: llm-model-pvc      # 모델 가중치는 이미지가 아닌 PVC에서 마운트
llm-inference-hpa.yamlYAML
# GPU Pod 오토스케일링 — CPU가 아닌 큐 길이/동시 요청 수 기반 커스텀 메트릭 권장
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: llm-inference-hpa
  namespace: ai-agent
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: llm-inference
  minReplicas: 1               # GPU는 고가 자원 — 최소치를 낮게 유지
  maxReplicas: 6               # 클러스터 내 확보 가능한 GPU 수로 상한 설정
  metrics:
  - type: Pods
    pods:
      metric:
        name: inference_queue_length   # Prometheus Adapter로 노출한 커스텀 메트릭
      target:
        type: AverageValue
        averageValue: "5"
구성 요소역할비고
NVIDIA device plugin (DaemonSet)GPU 노드를 nvidia.com/gpu 스케줄링 리소스로 노출GPU 노드에만 자동 배포됨
nodeSelector / taint-tolerationGPU Pod를 GPU 노드에만 스케줄, 일반 Pod는 배제GPU 노드에 NoSchedule taint를 걸어 두는 것이 일반적
모델 가중치 스토리지컨테이너 이미지에 넣지 않고 PVC/오브젝트 스토리지에서 마운트이미지 크기·배포 시간 단축, 모델 버전 교체 용이
Readiness probe모델 로딩이 끝나기 전 트래픽 유입 차단LLM은 로딩에 수십 초~수 분 소요 — initialDelaySeconds를 넉넉히

AI 모델 배포 전략 & 운영

여기서는 AI 모델 배포 전략 & 운영을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

LLM 추론 서버는 일반 웹 서비스와 운영 리스크가 다릅니다. GPU는 비싸서 최소 복제본을 낮게 유지해야 하고, 모델 하나를 바꾸는 것이 곧 수십 GB 가중치의 재검증을 의미하며, CPU/메모리 지표만으로는 지금 큐에 얼마나 많은 요청이 밀려 있는지 알 수 없습니다. 이 특성에 맞춰 점진적 모델 롤아웃, 큐 기반 오토스케일링, GPU 노드 장애 격리를 별도로 설계해야 합니다.
다이어그램 렌더링 중…
model-canary-route.yaml — Gateway API로 두 모델 버전에 트래픽 분할YAML
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: llm-inference-route
spec:
  parentRefs:
  - name: main-gateway
  rules:
  - backendRefs:
    - name: llm-inference-stable
      port: 8000
      weight: 95        # 검증된 기존 모델
    - name: llm-inference-canary
      port: 8000
      weight: 5         # 신규 모델 — 점진적으로 weight를 늘려감
keda-scaledobject.yaml — 큐 길이 기반 오토스케일링YAML
# Prometheus Adapter 없이도 Redis/SQS 등 외부 큐 지표로 바로 스케일링
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llm-inference-scaler
  namespace: ai-agent
spec:
  scaleTargetRef:
    name: llm-inference
  minReplicaCount: 1        # GPU는 고가 자원 — 유휴 상태에서는 0까지도 검토 가능
  maxReplicaCount: 6
  cooldownPeriod: 300        # 스케일 다운 전 대기 — GPU Pod는 축소·재확보 비용이 크므로 여유 있게
  triggers:
  - type: redis
    metadata:
      address: redis.ai-agent.svc:6379
      listName: inference-queue
      listLength: "5"        # 큐에 5개 이상 쌓이면 스케일 아웃
gpu-pdb.yaml — GPU Pod 동시 축출 방지YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: llm-inference-pdb
  namespace: ai-agent
spec:
  minAvailable: 1            # 노드 유지보수 중에도 최소 1개는 서비스 가능 상태 유지
  selector:
    matchLabels: { app: llm-inference }
운영 항목일반 웹 서비스LLM 추론 서버
배포 리스크이미지 하나 교체모델 가중치가 바뀌면 응답 품질 자체가 달라질 수 있음 — 카나리로 먼저 검증 필요
스케일링 신호CPU/메모리 사용률로 충분한 경우가 많음큐 길이·동시 요청 수 등 외부 지표 기반 스케일링이 더 정확
장애 시 복구 비용Pod 재시작 수 초GPU 노드 재확보·모델 재로딩까지 수 분 — PodDisruptionBudget으로 동시 축출 방지 필수

Tip

  • 카나리 모델을 실제 트래픽 5%로 검증할 때는 응답 지연시간뿐 아니라 토큰당 응답 품질(agent-evaluation 가이드의 LLM-as-Judge)과 토큰당 비용까지 함께 비교해야 "더 빠르지만 답변 품질이 나빠진 모델"을 놓치지 않습니다.
  • GPU 노드는 spot/preemptible 인스턴스를 쓰면 비용을 크게 줄일 수 있지만 예고 없이 회수될 수 있습니다 — PodDisruptionBudget과 함께, on-demand GPU 노드 풀을 fallback으로 두는 node affinity(preferredDuringScheduling)를 구성해두면 spot 회수 시에도 서비스가 완전히 끊기지 않습니다.

운영 & 업그레이드

여기서는 운영 & 업그레이드을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

kubectl 치트시트, 노드 유지보수, 클러스터 버전 업그레이드, etcd 백업·복구, 장애 진단 패턴을 정리합니다. 클러스터를 구축하는 것보다 몇 배 더 오래 지속되는 것이 일상 운영이므로, 노드를 안전하게 비우고 복귀시키는 절차와 etcd 백업처럼 장애 시 되돌릴 수 있는 수단을 미리 손에 익혀두는 것이 실제 사고 대응 속도를 좌우합니다.
kubectl-cheatsheet.shBASH
# ── kubectl 필수 명령어 ───────────────────────────────────────

# 컨텍스트 전환
kubectl config get-contexts
kubectl config use-context production

# 네임스페이스 기본 설정
kubectl config set-context --current --namespace=production

# Pod 상태 실시간 감시
kubectl get pods -n production -w
kubectl get pods -n production -o wide          # 노드 배치 포함

# 로그
kubectl logs <pod> -f                           # 실시간
kubectl logs <pod> --previous                   # 이전 컨테이너 로그
kubectl logs <pod> -c <container> --tail=100    # 멀티 컨테이너

# 디버깅
kubectl describe pod <pod> -n production
kubectl exec -it <pod> -- /bin/sh
kubectl debug <pod> -it --image=busybox         # 임시 디버그 컨테이너

# 포트 포워딩 (로컬 테스트)
kubectl port-forward svc/postgres-svc 5432:5432 -n database
kubectl port-forward pod/redis-0 6379:6379 -n cache

# 리소스 사용량
kubectl top nodes
kubectl top pods -n production --sort-by=cpu

# 이벤트 확인
kubectl get events -n production --sort-by='.lastTimestamp'

# 강제 재시작 (Deployment 롤링 재시작)
kubectl rollout restart deployment/<name> -n production

# 리소스 상세 출력
kubectl get all -n production
kubectl get all,pvc,ingress,cm,secret -n production
node-maintenance.shBASH
# ── 노드 유지보수 (cordon + drain) ───────────────────────────

# 1. 새 Pod 스케줄 차단 (기존 Pod 유지)
kubectl cordon k8s-worker-01

# 2. 기존 Pod를 다른 노드로 이동 (PodDisruptionBudget 존중)
kubectl drain k8s-worker-01 \
  --ignore-daemonsets \    # DaemonSet Pod는 드레인 제외
  --delete-emptydir-data \  # emptyDir 볼륨 Pod 삭제 허용
  --grace-period=30 \      # 종료 대기 시간
  --timeout=300s            # 전체 타임아웃

# 3. OS 패치, 커널 업그레이드, 디스크 교체 등 유지보수 수행

# 4. 노드 복구 후 스케줄 재개
kubectl uncordon k8s-worker-01

# 5. 노드 상태 확인
kubectl get node k8s-worker-01 -o wide
kubectl describe node k8s-worker-01 | grep -A5 "Conditions:"
cluster-upgrade.shBASH
# ── 클러스터 버전 업그레이드 (1.35 → 1.36, 예시) ──────────────────
# kubeadm은 minor 버전을 한 단계씩만 건너뛸 수 있음 — 1.34에서 1.36으로 직행 불가, 1.35를 반드시 경유
# 반드시 Control Plane → Worker 순서

# 0. 사전 체크
kubeadm upgrade plan

# 1. Control Plane 업그레이드
# kubeadm 업데이트
apt-mark unhold kubeadm
apt-get install -y kubeadm=1.36.0-00
apt-mark hold kubeadm

# 업그레이드 실행
kubeadm upgrade apply v1.36.0

# kubelet, kubectl 업데이트
apt-mark unhold kubelet kubectl
apt-get install -y kubelet=1.36.0-00 kubectl=1.36.0-00
apt-mark hold kubelet kubectl
systemctl daemon-reload
systemctl restart kubelet

# 2. Worker 노드 업그레이드 (각 노드 반복)
# Control Plane에서
kubectl cordon k8s-worker-01
kubectl drain k8s-worker-01 --ignore-daemonsets --delete-emptydir-data

# Worker 노드에서
apt-mark unhold kubeadm kubelet kubectl
apt-get install -y kubeadm=1.36.0-00 kubelet=1.36.0-00 kubectl=1.36.0-00
apt-mark hold kubeadm kubelet kubectl
kubeadm upgrade node
systemctl daemon-reload && systemctl restart kubelet

# Control Plane에서
kubectl uncordon k8s-worker-01

# 3. 업그레이드 완료 확인
kubectl get nodes
etcd-backup.shBASH
# ── etcd 백업 & 복구 ─────────────────────────────────────────
# etcd는 클러스터의 모든 상태 저장소 — 정기 백업 필수

# 백업 (Control Plane에서)
ETCD_BACKUP_DIR=/backup/etcd/$(date +%Y%m%d-%H%M%S)
mkdir -p $ETCD_BACKUP_DIR

ETCDCTL_API=3 etcdctl snapshot save $ETCD_BACKUP_DIR/snapshot.db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# 백업 검증
ETCDCTL_API=3 etcdctl snapshot status $ETCD_BACKUP_DIR/snapshot.db \
  --write-out=table

# cron 정기 백업 (매일 새벽 2시)
echo "0 2 * * * root ETCDCTL_API=3 etcdctl snapshot save /backup/etcd/$(date +%Y%m%d).db   --endpoints=https://127.0.0.1:2379   --cacert=/etc/kubernetes/pki/etcd/ca.crt   --cert=/etc/kubernetes/pki/etcd/server.crt   --key=/etc/kubernetes/pki/etcd/server.key" >> /etc/crontab

# ── 복구 절차 ────────────────────────────────────────────────
# 1. 모든 Control Plane에서 kubelet 중지
systemctl stop kubelet

# 2. etcd 데이터 디렉터리 백업 후 삭제
mv /var/lib/etcd /var/lib/etcd.bak

# 3. 스냅샷 복구
ETCDCTL_API=3 etcdctl snapshot restore /backup/etcd/20240101.db \
  --data-dir=/var/lib/etcd \
  --name=k8s-master-01 \
  --initial-cluster=k8s-master-01=https://192.168.1.10:2380 \
  --initial-cluster-token=etcd-cluster-1 \
  --initial-advertise-peer-urls=https://192.168.1.10:2380

# 4. kubelet 재시작
systemctl start kubelet
kubectl get nodes
troubleshoot.shBASH
# ── 장애 진단 패턴 ───────────────────────────────────────────

# Pod가 Pending 상태일 때
kubectl describe pod <pod> -n <ns>
# → Events 항목 확인:
#   "Insufficient cpu"    → ResourceQuota 또는 노드 자원 부족
#   "No nodes available"  → 모든 노드 taint/cordon 상태
#   "PVC not found"       → StorageClass 또는 PV 문제

# Pod가 CrashLoopBackOff일 때
kubectl logs <pod> --previous -n <ns>
kubectl describe pod <pod> -n <ns>
# → 종료 코드 확인:
#   Exit Code 1: 앱 오류 (로그 확인)
#   Exit Code 137: OOM Kill → limits.memory 증가

# 노드 NotReady일 때
kubectl describe node k8s-worker-01
# → Conditions: MemoryPressure, DiskPressure, PIDPressure 확인
ssh k8s-worker-01
systemctl status kubelet
journalctl -u kubelet -f

# 서비스 접근 불가일 때
kubectl get endpoints <service> -n <ns>    # Pod IP 등록 여부
kubectl exec <pod> -- curl http://<service>.<ns>.svc.cluster.local
kubectl exec <pod> -- nslookup <service>   # DNS 확인

# 이미지 Pull 실패
kubectl describe pod <pod> | grep -A3 "Failed"
# → ErrImagePull / ImagePullBackOff
#   private registry: imagePullSecrets 설정 확인
kubectl create secret docker-registry regcred \
  --docker-server=registry.aidevops.kr \
  --docker-username=<user> \
  --docker-password=<pass>

# 클러스터 전체 헬스 체크 스크립트
echo "=== Nodes ===" && kubectl get nodes
echo "=== System Pods ===" && kubectl get pods -n kube-system | grep -v Running
echo "=== Failed Pods ===" && kubectl get pods -A | grep -E "Error|CrashLoop|Pending|OOMKilled"
echo "=== PVC ===" && kubectl get pvc -A | grep -v Bound
echo "=== Recent Events ===" && kubectl get events -A --sort-by='.lastTimestamp' | tail -20

Helm 차트

여기서는 Helm 차트을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Helm은 여러 YAML 매니페스트를 하나의 배포 단위(차트)로 묶고, 환경마다 달라지는 값(이미지 태그, 레플리카 수, GPU 개수 등)만 values 파일로 바꿔 재사용할 수 있게 해줍니다. kubectl apply를 반복하는 것과 근본적으로 다른 점은 리비전 히스토리입니다 — helm upgrade를 실행할 때마다 그 시점의 매니페스트와 values 전체가 스냅샷으로 기록되어, helm rollback 한 줄로 이전 리비전 전체 상태로 되돌릴 수 있습니다.
다이어그램 렌더링 중…
차트 디렉터리 구조TEXT
llm-inference-chart/
├── Chart.yaml           # 차트 이름, 버전, 의존성(서브차트) 선언
├── values.yaml          # 기본값 — 모든 환경의 출발점
├── values-staging.yaml   # 스테이징 전용 override (레플리카 1, GPU 1개)
├── values-production.yaml  # 운영 전용 override (레플리카 3, GPU 타입 지정)
├── templates/
│   ├── deployment.yaml   # {{ .Values.* }} 로 값이 채워지는 템플릿
│   ├── service.yaml
│   ├── hpa.yaml
│   ├── _helpers.tpl      # 여러 템플릿에서 재사용할 이름·라벨 헬퍼 함수
│   └── hooks/
│       └── pre-upgrade-migration.yaml
└── charts/               # 의존 서브차트 (예: postgresql) — helm dependency update로 채워짐
templates/deployment.yaml — gpu-inference-deploy 섹션의 Deployment를 템플릿화YAML
apiVersion: apps/v1
kind: Deployment
metadata:
  name: {{ include "llm-inference.fullname" . }}
  labels:
    {{- include "llm-inference.labels" . | nindent 4 }}
spec:
  replicas: {{ .Values.replicaCount }}
  selector:
    matchLabels:
      app: {{ include "llm-inference.fullname" . }}
  template:
    metadata:
      labels:
        app: {{ include "llm-inference.fullname" . }}
    spec:
      nodeSelector:
        gpu-type: {{ .Values.gpu.nodeType }}
      containers:
      - name: vllm
        image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
        args: ["--model", "{{ .Values.model.path }}"]
        resources:
          requests:
            nvidia.com/gpu: {{ .Values.gpu.count }}
          limits:
            nvidia.com/gpu: {{ .Values.gpu.count }}
values-production.yaml — 환경별 overrideYAML
replicaCount: 3
image:
  repository: vllm/vllm-openai
  tag: "0.6.2"
gpu:
  nodeType: a100    # 스테이징은 values-staging.yaml에서 a10g로 override
  count: 1
model:
  path: /models/llama-3-70b
templates/hooks/pre-upgrade-migration.yaml — 업그레이드 전 자동 실행되는 JobYAML
apiVersion: batch/v1
kind: Job
metadata:
  name: {{ include "llm-inference.fullname" . }}-migrate
  annotations:
    "helm.sh/hook": pre-upgrade,pre-install
    "helm.sh/hook-delete-policy": hook-succeeded   # 성공한 Job은 정리, 실패 시 원인 조사를 위해 남김
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: migrate
        image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
        command: ["python", "scripts/sync_model_registry.py"]
BASH
# 설치 & 배포
helm create myapp                       # 기본 차트 스캐폴딩 생성
helm install myapp ./myapp -f values-production.yaml -n production --atomic --wait
#   --atomic: 실패 시 자동으로 이전 상태로 롤백
#   --wait:   Pod가 Ready 상태가 될 때까지 명령이 종료되지 않고 대기

# 실제 적용 전에 렌더링 결과 확인 (클러스터에 아무것도 적용하지 않음)
helm template myapp ./myapp -f values-production.yaml

# 업그레이드 & 롤백
helm upgrade myapp ./myapp -f values-production.yaml --set image.tag=1.3.0 --atomic
helm history myapp                       # 리비전별 변경 요약
helm rollback myapp 2                    # 특정 리비전으로 즉시 복귀

# 서브차트(postgresql 등) 의존성 다운로드
helm dependency update ./myapp

# 차트를 OCI 레지스트리에 패키징 & 배포 (최신 Helm 표준 배포 방식)
helm package ./myapp
helm push myapp-1.0.0.tgz oci://ghcr.io/myorg/charts
명령어/플래그역할
helm template클러스터에 아무것도 적용하지 않고 렌더링된 YAML만 미리 확인 (dry-run)
--atomic업그레이드 실패 시 자동으로 이전 리비전으로 롤백 — 운영 배포에는 거의 필수
--waitPod가 실제로 Ready 상태가 될 때까지 대기 후 명령 종료 — CI 파이프라인에서 배포 성공 여부를 정확히 판단
helm.sh/hook (pre-upgrade 등)DB 마이그레이션처럼 배포 전후에 한 번만 실행해야 하는 작업을 차트에 포함
charts/ (서브차트)postgresql, redis 같은 의존 컴포넌트를 다른 팀이 만든 차트 그대로 가져와 조합 (umbrella chart)

Tip

values.yaml에 DB 비밀번호 같은 시크릿을 평문으로 넣지 마세요 — 그 값도 helm history에 리비전마다 그대로 남습니다. 시크릿은 values 밖에서 External Secrets Operator 등으로 주입하고, 차트에는 Secret 리소스의 이름(참조)만 넣으세요.

Contour & Gateway API — NGINX Ingress 대체

여기서는 Contour & Gateway API — NGINX Ingress 대체을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Contour는 Envoy Proxy를 데이터 플레인으로 쓰는 L7 Ingress Controller입니다. NGINX Ingress와 달리 설정 변경 시 워커 프로세스 reload가 필요 없이 xDS API로 Envoy 설정을 무중단으로 갱신하며, 대규모 라우트에서도 reload로 인한 커넥션 드롭이 없습니다. 라우팅 정책은 두 가지 방식으로 선언할 수 있습니다 — Contour 전용 CRD인 HTTPProxy(팀 간 위임·세밀한 정책에 강점)와, 컨트롤러 벤더에 종속되지 않는 Kubernetes 표준 Gateway API(GatewayClass/Gateway/HTTPRoute 3계층 모델)입니다. 신규 구축이면 Gateway API를, 팀별 네임스페이스 위임이나 세밀한 IP/Rate limit 정책이 필요하면 HTTPProxy를 우선 검토하세요.
contour-install.shBASH
# 방법 A — HTTPProxy 중심 (정적 매니페스트, IngressClass "contour" 생성)
kubectl apply -f https://projectcontour.io/quickstart/contour.yaml
kubectl get pods -n projectcontour -w
kubectl get svc envoy -n projectcontour   # LoadBalancer 외부 IP 확인

# 방법 B — Gateway API 중심 (Contour Gateway Provisioner)
# 1) Gateway API CRD 설치 (standard channel)
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.1.0/standard-install.yaml

# 2) Contour + Gateway Provisioner 설치 — GatewayClass "contour" 자동 생성
kubectl apply -f https://projectcontour.io/quickstart/contour-gateway-provisioner.yaml
kubectl get gatewayclass contour

# 이후 Gateway 리소스를 생성하면 Provisioner가 해당 Gateway 전용 Contour+Envoy 워크로드를 자동 배포합니다.
gateway.yamlYAML
# GatewayClass는 contour-gateway-provisioner.yaml 적용 시 자동 생성됨
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: contour-gateway
  namespace: projectcontour
spec:
  gatewayClassName: contour
  listeners:
  - name: http
    protocol: HTTP
    port: 80
    hostname: "*.aidevops.kr"
    allowedRoutes:
      namespaces:
        from: All
  - name: https
    protocol: HTTPS
    port: 443
    hostname: "*.aidevops.kr"
    tls:
      mode: Terminate
      certificateRefs:
      - kind: Secret
        name: wildcard-aidevops-kr-tls   # cert-manager가 DNS-01로 발급
    allowedRoutes:
      namespaces:
        from: All
httproute-canary.yamlYAML
# Gateway API HTTPRoute — 헤더 기반 라우팅 + 가중치 카나리 배포
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api-route
  namespace: production
spec:
  parentRefs:
  - name: contour-gateway
    namespace: projectcontour
  hostnames:
  - "api.aidevops.kr"
  rules:
  # x-canary: true 헤더가 있으면 카나리 서비스로 100% 라우팅
  - matches:
    - headers:
      - name: x-canary
        value: "true"
    backendRefs:
    - name: api-service-canary
      port: 8080
  # 나머지 트래픽은 90:10 가중치 분산
  - backendRefs:
    - name: api-service
      port: 8080
      weight: 90
    - name: api-service-canary
      port: 8080
      weight: 10
httpproxy-policy.yamlYAML
# Contour HTTPProxy — 루트: 도메인/TLS/전역 Rate limit 정의 + 팀별 경로 위임(includes)
apiVersion: projectcontour.io/v1
kind: HTTPProxy
metadata:
  name: aidevops-root
  namespace: production
spec:
  virtualhost:
    fqdn: api.aidevops.kr
    tls:
      secretName: wildcard-aidevops-kr-tls
    rateLimitPolicy:
      local:
        requests: 100
        unit: second
        burst: 50
  includes:
  - name: api-v1-routes          # production 네임스페이스 팀이 관리
    namespace: production
    conditions:
    - prefix: /api/v1
  - name: api-v2-routes          # payments 팀이 자신의 네임스페이스에서 독립적으로 관리
    namespace: team-payments
    conditions:
    - prefix: /api/v2
  routes:
  - conditions:
    - prefix: /
    services:
    - name: frontend-service
      port: 3000
---
# production 네임스페이스에서 위임받은 자식 HTTPProxy — 세부 정책은 팀이 직접 소유
apiVersion: projectcontour.io/v1
kind: HTTPProxy
metadata:
  name: api-v1-routes
  namespace: production
spec:
  routes:
  - conditions:
    - prefix: /api/v1
    services:
    - name: api-service
      port: 8080
      weight: 90
    - name: api-service-canary
      port: 8080
      weight: 10
    retryPolicy:
      count: 3
      perTryTimeout: 2s
    timeoutPolicy:
      response: 10s
      idle: 60s
    ipAllowFilterPolicy:         # 사내망 + 지정 대역만 허용
    - source: 10.0.0.0/8
    - source: 203.0.113.0/24
wildcard-dns01-clusterissuer.yamlYAML
# 와일드카드 도메인(*.aidevops.kr)은 HTTP-01로 발급 불가 — DNS-01 challenge 필수
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-dns01
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: admin@aidevops.kr
    privateKeySecretRef:
      name: letsencrypt-dns01
    solvers:
    - dns01:
        route53:
          region: ap-northeast-2
          hostedZoneID: ZXXXXXXXXXXXXX   # Route53 Hosted Zone ID
      selector:
        dnsZones:
        - "aidevops.kr"
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: wildcard-aidevops-kr
  namespace: projectcontour
spec:
  secretName: wildcard-aidevops-kr-tls
  issuerRef:
    name: letsencrypt-dns01
    kind: ClusterIssuer
  dnsNames:
  - "aidevops.kr"
  - "*.aidevops.kr"
항목NGINX IngressContour HTTPProxyGateway API HTTPRoute
설정 반영reload 필요 — 대량 Ingress에서 커넥션 드롭 위험Envoy xDS로 무중단 동적 반영동일 (Contour가 Envoy로 xDS 반영)
팀 간 위임annotation 기반, 네임스페이스 간 위임 어려움`spec.includes`로 경로/도메인별 위임 명시ReferenceGrant로 크로스 네임스페이스 허용
카나리/가중치 라우팅벤더별 annotation(canary-weight)`services[].weight` 필드`backendRefs[].weight` 필드
헤더 기반 라우팅annotation 조합 필요`conditions[].header``matches[].headers`
표준 여부컨트롤러마다 비표준 annotationContour 전용 CRDKubernetes SIG 공식 표준(k8s.io)
Rate limit / IP 정책컨트롤러마다 상이한 annotation`rateLimitPolicy`, `ipAllow/DenyFilterPolicy`GEP 확장 또는 별도 정책 컨트롤러 필요

Proxy Protocol — LB 뒤에서 클라이언트 IP 보존하기

여기서는 Proxy Protocol — LB 뒤에서 클라이언트 IP 보존하기을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Contour(Envoy) 앞에 L4 LoadBalancer를 두면 Envoy 액세스 로그나 애플리케이션이 보는 클라이언트 IP가 실제 사용자 IP가 아니라 LB나 노드의 IP로 찍히는 문제가 자주 생깁니다. LB가 SNAT을 하거나(AWS NLB의 ip 타깃 모드, Classic ELB, 여러 매니지드 LB), kube-proxy가 externalTrafficPolicy: Cluster에서 노드 간 전달 시 SNAT을 하기 때문입니다. HTTPS는 Envoy가 TLS를 종료하므로 L4 LB는 암호화된 바이트를 들여다볼 수 없어 X-Forwarded-For 헤더를 넣어줄 수도 없습니다.

Proxy Protocol은 HAProxy가 제안한 규약으로, L4 LB가 백엔드로 TCP 연결을 맺자마자 연결 맨 앞에 원본 클라이언트 IP/포트를 담은 짧은 헤더를 먼저 보내고, 그 뒤에 원래의 HTTP/TLS 바이트를 그대로 흘려보내는 방식입니다. v1은 사람이 읽을 수 있는 텍스트 한 줄, v2는 바이너리 형식이며 v2는 TLV 확장(AWS VPC Endpoint ID 등)을 실을 수 있습니다. Envoy의 envoy.filters.listener.proxy_protocol 리스너 필터는 v1/v2를 자동 인식해 헤더를 읽어낸 뒤, 이후 요청의 downstream remote address를 헤더 속 원본 IP로 바꿉니다 — 그래서 Envoy가 업스트림으로 보내는 X-Forwarded-For, X-Envoy-External-Address, 액세스 로그, ipAllowFilterPolicy 판정이 모두 실제 클라이언트 IP 기준으로 동작합니다.

⚠️
핵심 규칙 — 보내는 쪽과 받는 쪽이 반드시 동시에 켜져야 합니다
Proxy Protocol 스펙은 수신 측이 헤더 유무를 "추측"하지 말고, 헤더가 필요한 리스너에서는 헤더 없는 연결을 거부하도록 정의합니다. Contour에서 Proxy Protocol을 켜면 Envoy의 모든 트래픽 리스너(HTTP 8080, HTTPS 8443)에 이 필터가 붙기 때문에, LB를 거치지 않고 Envoy로 직접 들어오는 클러스터 내부 호출은 전부 끊깁니다. 반대로 LB만 켜고 Envoy를 켜지 않으면 헤더가 HTTP/TLS 데이터로 해석되어 400 Bad Request나 TLS 핸드셰이크 실패가 납니다. 이 가이드의 이후 하위 섹션은 이 두 가지 문제를 LB·Contour 버전·대체 방법별로 풀어갑니다.
다이어그램 렌더링 중…
proxy-protocol-header-example.txtTEXT
# v1 (텍스트) — 연결 직후 한 줄이 먼저 전송되고, 그 뒤에 원래 HTTP/TLS 데이터가 이어짐
PROXY TCP4 203.0.113.7 10.0.12.34 51234 443\r\n
#     프로토콜 원본IP        LB/수신IP  원본포트 목적지포트

# v2 (바이너리) — 12바이트 시그니처로 시작
\x0D\x0A\x0D\x0A\x00\x0D\x0A\x51\x55\x49\x54\x0A  + 버전/명령 + 주소 + TLV(선택)

# Envoy 리스너 필터 구성 (Contour가 useProxyProtocol=true일 때 생성하는 형태)
listener_filters:
- name: envoy.filters.listener.proxy_protocol
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.listener.proxy_protocol.v3.ProxyProtocol
    # allow_requests_without_proxy_protocol: false  ← 기본값. Contour는 이 값을 바꿀 옵션을 제공하지 않음
- name: envoy.filters.listener.tls_inspector    # HTTPS 리스너는 SNI 확인을 위해 뒤이어 실행
클라이언트 IP 보존 방식적합한 환경장점주의점
externalTrafficPolicy: LocalLB가 SNAT 없이 패킷을 그대로 전달(패스스루)하는 경우 — GKE/AKS L4 LB, MetalLB, AWS NLB instance 타깃Service 설정 한 줄, 내부 호출에 영향 없음Envoy Pod가 없는 노드는 LB 헬스체크에서 빠짐 → 노드별 트래픽 불균형 가능
Proxy ProtocolLB가 SNAT/프록시를 하는 경우 — AWS NLB ip 타깃, Classic ELB, NCP, DigitalOcean, HAProxy/F5 등TLS를 Envoy가 종료해도 원본 IP 전달, externalTrafficPolicy와 무관LB·Envoy 동시 설정 필요, 헤더 없는 내부 호출 차단, 헤어핀 트래픽 문제
L7 LB + X-Forwarded-ForALB, Application Gateway 등 L7 LB가 TLS를 먼저 종료하는 경우Proxy Protocol 불필요, WAF 연동 쉬움Contour num-trusted-hops 설정 필수, TLS가 두 번 종료됨

Tip

  • Proxy Protocol은 L4 기술입니다 — HTTP 헤더가 아니라 TCP 스트림 맨 앞의 바이트이므로, 애플리케이션 코드는 수정할 필요 없이 Envoy가 만들어주는 X-Forwarded-For만 읽으면 됩니다.
  • LB가 패스스루 방식이라면 Proxy Protocol보다 externalTrafficPolicy: Local이 훨씬 단순합니다. 먼저 사용 중인 LB가 SNAT을 하는지부터 확인하세요.
  • Proxy Protocol을 켜는 순간 "LB를 거치지 않는 모든 경로"가 영향을 받습니다 — 적용 전에 클러스터 내부에서 Envoy Service나 외부 도메인을 호출하는 워크로드가 있는지 먼저 목록화하세요.

LB별 Proxy Protocol 설정

여기서는 LB별 Proxy Protocol 설정을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Proxy Protocol 헤더를 붙이는 주체는 LB입니다. Kubernetes에서는 Envoy를 노출하는 type: LoadBalancer Service에 클라우드별 annotation을 달아 Cloud Controller Manager(또는 AWS Load Balancer Controller)가 LB 리스너/타깃 그룹에 Proxy Protocol을 켜도록 합니다. Proxy Protocol을 쓰면 원본 IP가 헤더로 전달되므로 externalTrafficPolicy는 Cluster로 두어도 되고, 노드 간 트래픽 균형 측면에서는 오히려 유리합니다.

헬스체크 주의: LB 헬스체크가 트래픽 포트로 연결할 때 PROXY 헤더를 보내지 않거나, 반대로 헤더를 이해하지 못하는 포트(Envoy의 8002 /ready, kube-proxy healthCheckNodePort)로 헤더를 보내면 타깃이 unhealthy로 빠집니다. 클라우드 문서에서 "헬스체크에도 Proxy Protocol 헤더가 포함되는지"를 확인하고, 애매하면 트래픽 포트에 대한 TCP 헬스체크로 시작해 동작을 확인한 뒤 HTTP 헬스체크로 옮기세요.

Envoy의 리스너 필터는 v1과 v2를 모두 자동 인식하므로 LB가 어떤 버전을 보내든 Contour 쪽 설정은 동일합니다.
envoy-service-aws-nlb.yamlYAML
# 정적 매니페스트(contour.yaml)로 설치한 경우 — envoy Service에 annotation 추가
# Gateway Provisioner를 쓴다면 이 Service는 Provisioner가 관리하므로 직접 수정하지 말고
# ContourDeployment.spec.envoy.networkPublishing.serviceAnnotations로 지정 (다음 섹션)
apiVersion: v1
kind: Service
metadata:
  name: envoy
  namespace: projectcontour
  annotations:
    service.beta.kubernetes.io/aws-load-balancer-type: external
    service.beta.kubernetes.io/aws-load-balancer-scheme: internet-facing
    service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
    service.beta.kubernetes.io/aws-load-balancer-proxy-protocol: "*"   # 모든 타깃 그룹에 PP v2
spec:
  type: LoadBalancer
  externalTrafficPolicy: Cluster     # 원본 IP는 PROXY 헤더로 전달되므로 Cluster로도 충분
  selector:
    app: envoy
  ports:
  - name: http
    port: 80
    targetPort: 8080
    protocol: TCP
  - name: https
    port: 443
    targetPort: 8443
    protocol: TCP
aws-verify-target-group.shBASH
# NLB 타깃 그룹에 Proxy Protocol v2가 실제로 켜졌는지 확인
aws elbv2 describe-target-groups \
  --load-balancer-arn "$NLB_ARN" \
  --query "TargetGroups[].TargetGroupArn" --output text |
tr '\t' '\n' | while read -r TG; do
  echo "== $TG"
  aws elbv2 describe-target-group-attributes --target-group-arn "$TG" \
    --query "Attributes[?Key=='proxy_protocol_v2.enabled']"
done
haproxy.cfg (온프레미스 L4 앞단)TEXT
frontend fe_https
    mode tcp
    bind :443
    default_backend be_envoy_https

backend be_envoy_https
    mode tcp
    balance roundrobin
    # send-proxy-v2: 백엔드(Envoy NodePort)로 연결 시 PROXY v2 헤더 전송
    # check-send-proxy: 헬스체크 연결에도 헤더를 붙여 Envoy 리스너가 거부하지 않도록 함
    server node1 10.0.0.11:30443 send-proxy-v2 check check-send-proxy
    server node2 10.0.0.12:30443 send-proxy-v2 check check-send-proxy
nginx-stream.conf (온프레미스 L4 앞단)NGINX
stream {
    upstream envoy_https {
        server 10.0.0.11:30443;
        server 10.0.0.12:30443;
    }
    server {
        listen 443;
        proxy_pass envoy_https;
        proxy_protocol on;     # 업스트림(Envoy)으로 PROXY v1 헤더 전송
    }
}
LBProxy Protocol 활성화 방법버전비고
AWS NLB (AWS Load Balancer Controller)service.beta.kubernetes.io/aws-load-balancer-proxy-protocol: "*"v2ip 타깃 모드는 클라이언트 IP 보존이 기본 비활성 → Proxy Protocol이 사실상 필수. 타깃 그룹 속성 proxy_protocol_v2.enabled로 확인
AWS Classic ELB (in-tree, 레거시)위와 같은 annotationv1신규 구축에는 NLB 권장
GKE L4 (패스스루 NLB)해당 없음 — externalTrafficPolicy: Local 사용-패킷을 그대로 전달하므로 원본 IP가 이미 보존됨
Azure LB (AKS)해당 없음 — externalTrafficPolicy: Local 사용-Private Link Service 경유 시에만 service.beta.kubernetes.io/azure-pls-proxy-protocol: "true"
NCP NKS (Network Proxy LB)service.beta.kubernetes.io/ncloud-load-balancer-proxy-protocol: "true"v2LB 타입·NKS 버전별 지원 annotation을 NCP 문서에서 확인
DigitalOceanservice.beta.kubernetes.io/do-loadbalancer-enable-proxy-protocol: "true"v1LB가 IP를 반환하므로 헤어핀 문제 대응 필요 (내부 호출 섹션 참고)
온프레미스 HAProxy / F5 / NGINX streamHAProxy send-proxy-v2, NGINX proxy_protocol on;, F5 iRulev1/v2MetalLB(L2/BGP)는 Proxy Protocol을 붙이지 않음 → Local 정책 사용

Tip

  • LB annotation 변경은 클라우드에 따라 LB 재생성을 유발할 수 있습니다(외부 IP/DNS 변경). 운영 중인 Service라면 새 Service를 병행 생성해 DNS를 전환하는 방식이 안전합니다.
  • LB와 Envoy 설정을 "동시에" 바꿀 수 없으므로 전환 순간 짧은 장애가 생깁니다. 무중단이 필요하면 Proxy Protocol이 켜진 새 LB + 새 Envoy 세트를 만들고 DNS 가중치로 넘기세요.

Contour 버전별 Envoy Proxy Protocol 설정 (1.28 ~ 1.33)

여기서는 Contour 버전별 Envoy Proxy Protocol 설정 (1.28 ~ 1.33)을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Contour에서 Envoy가 PROXY 헤더를 받도록 하는 방법은 설치 방식에 따라 세 가지입니다. 어느 방식이든 결과는 같습니다 — Contour가 Envoy의 ingress_http/ingress_https 리스너 listener_filters 맨 앞에 envoy.filters.listener.proxy_protocol을 추가합니다. 통계용 리스너(8002)와 admin 인터페이스에는 적용되지 않습니다.
설치 방식설정 위치필드/플래그
정적 매니페스트 (quickstart contour.yaml)contour Deployment의 contour serve 인자--use-proxy-protocol
ContourConfiguration CRD 사용 (contour serve --contour-config-name)ContourConfigurationspec.envoy.listener.useProxyProtocol: true
Gateway Provisioner (Gateway API)ContourDeployment (GatewayClass의 parametersRef)spec.runtimeSettings.envoy.listener.useProxyProtocol: true
주의: Contour 설정 파일(contour ConfigMap의 contour.yaml)에는 Proxy Protocol에 대응하는 키가 없습니다. ConfigMap에 넣어도 무시되므로 플래그나 CRD를 사용하세요.

아래 버전 표는 Contour 공식 compatibility matrix 기준입니다. 1.28 ~ 1.33 사이에서 Proxy Protocol 설정 방식 자체는 바뀌지 않았고, 달라지는 것은 번들 Envoy 버전과 지원 Kubernetes·Gateway API 버전입니다. 특히 1.33은 패치 릴리스 도중(1.33.6) 번들 Envoy가 1.35 → 1.38로 크게 올라갔으므로, 1.33.x 안에서의 패치 업그레이드라도 Envoy 릴리스 노트의 동작 변경 사항을 확인하고 스테이징에서 먼저 검증하세요.
A-static-manifest-flag.shBASH
# [방법 A] 정적 매니페스트(quickstart contour.yaml) — contour serve에 플래그 추가
kubectl -n projectcontour patch deployment contour --type=json -p='[
  {"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--use-proxy-protocol"}
]'

kubectl -n projectcontour rollout status deployment/contour
kubectl -n projectcontour get deploy contour \
  -o jsonpath='{.spec.template.spec.containers[0].args}' | tr ',' '\n'
# ["serve", "--incluster", ..., "--use-proxy-protocol"]

# ※ quickstart 매니페스트를 다시 apply하면 플래그가 사라집니다.
#   Kustomize patch 또는 GitOps 저장소에 패치를 함께 관리하세요.
B-contourconfiguration.yamlYAML
# [방법 B] ContourConfiguration CRD — contour serve --contour-config-name=contour 로 실행할 때 사용
apiVersion: projectcontour.io/v1alpha1
kind: ContourConfiguration
metadata:
  name: contour
  namespace: projectcontour
spec:
  envoy:
    listener:
      useProxyProtocol: true       # 모든 Envoy 트래픽 리스너에 proxy_protocol 리스너 필터 추가
    network:
      numTrustedHops: 0            # 앞단이 L4(PP)만 있으면 0 — L7 프록시가 한 단계 더 있으면 그 수만큼
  # ... 기존 설정(gateway, httpproxy, tls 등)은 그대로 유지
C-contourdeployment-provisioner.yamlYAML
# [방법 C] Gateway Provisioner — GatewayClass별로 Contour+Envoy 설정을 분리 관리
apiVersion: projectcontour.io/v1alpha1
kind: ContourDeployment
metadata:
  name: contour-external-pp
  namespace: projectcontour
spec:
  runtimeSettings:                 # = ContourConfiguration spec과 동일한 구조
    envoy:
      listener:
        useProxyProtocol: true
  envoy:
    networkPublishing:
      type: LoadBalancerService
      externalTrafficPolicy: Cluster
      serviceAnnotations:          # Provisioner가 만드는 envoy Service에 그대로 전달됨
        service.beta.kubernetes.io/aws-load-balancer-type: external
        service.beta.kubernetes.io/aws-load-balancer-scheme: internet-facing
        service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
        service.beta.kubernetes.io/aws-load-balancer-proxy-protocol: "*"
---
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
  name: contour-external-pp
spec:
  controllerName: projectcontour.io/gateway-controller
  parametersRef:
    group: projectcontour.io
    kind: ContourDeployment
    name: contour-external-pp
    namespace: projectcontour
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: external-gateway
  namespace: projectcontour
spec:
  gatewayClassName: contour-external-pp
  listeners:
  - name: http
    protocol: HTTP
    port: 80
    allowedRoutes:
      namespaces:
        from: All
  - name: https
    protocol: HTTPS
    port: 443
    hostname: "*.aidevops.kr"
    tls:
      mode: Terminate
      certificateRefs:
      - kind: Secret
        name: wildcard-aidevops-kr-tls
    allowedRoutes:
      namespaces:
        from: All
D-helm-values.yamlYAML
# [참고] Helm 차트로 설치한 경우 — 차트마다 키 이름이 다르므로 반드시 확인:
#   helm show values <repo>/contour --version <차트버전> | grep -n -i -E "extraArgs|proxy|annotations"
contour:
  extraArgs:
  - --use-proxy-protocol
envoy:
  service:
    externalTrafficPolicy: Cluster
    annotations:
      service.beta.kubernetes.io/aws-load-balancer-type: external
      service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
      service.beta.kubernetes.io/aws-load-balancer-proxy-protocol: "*"
check-contour-envoy-version.shBASH
# 현재 클러스터의 Contour / Envoy 이미지 버전 확인 → 위 표에서 해당 행 찾기
kubectl -n projectcontour get deploy,ds \
  -o custom-columns='NAME:.metadata.name,IMAGES:.spec.template.spec.containers[*].image'
# NAME      IMAGES
# contour   ghcr.io/projectcontour/contour:v1.33.x
# envoy     ghcr.io/projectcontour/contour:v1.33.x,docker.io/envoyproxy/envoy:v1.35.x

# 설정이 실제 Envoy 리스너에 반영됐는지 확인 (Envoy admin: Pod 내부 9001)
ENVOY_POD=$(kubectl -n projectcontour get pod -l app=envoy -o jsonpath='{.items[0].metadata.name}')
kubectl -n projectcontour port-forward "$ENVOY_POD" 9001:9001 >/dev/null &
sleep 2
curl -s localhost:9001/config_dump | grep -c "envoy.filters.listener.proxy_protocol"
# 2 이상 (ingress_http + ingress_https) 이면 적용됨, 0이면 미적용
# ※ Provisioner로 만든 Envoy는 라벨이 다를 수 있음: kubectl get pod -n projectcontour --show-labels
Contour번들 EnvoyKubernetesGateway APIProxy Protocol 설정PROXY 없는 연결 허용 (allow_requests_without_proxy_protocol)
1.28.x1.29.x1.27 ~ 1.29v1.0.0플래그 / ContourConfiguration / ContourDeploymentEnvoy는 지원, Contour 미노출 → 설정 불가
1.29.x1.30.x1.27 ~ 1.29v1.0.0동일설정 불가
1.30.x1.31.x1.28 ~ 1.30v1.1.0동일설정 불가
1.31.x1.34.x1.30 ~ 1.32v1.2.1동일설정 불가
1.32.x1.34.x1.31 ~ 1.33v1.2.1동일설정 불가
1.33.0 ~ 1.33.51.35.x1.32 ~ 1.34v1.3.0동일설정 불가
1.33.6 ~1.38.x1.32 ~ 1.34v1.3.0동일설정 불가 — 대체 방법은 다음 섹션

Tip

  • Proxy Protocol 설정은 Contour 전체(모든 Envoy 트래픽 리스너)에 일괄 적용됩니다. 리스너별로 켜고 끌 수 없으므로 "외부용(PP on)"과 "내부용(PP off)"을 나누려면 Contour+Envoy 세트를 두 개 운영해야 합니다.
  • Gateway Provisioner를 쓰면 GatewayClass(ContourDeployment)마다 Contour+Envoy 세트가 따로 생기므로, 외부/내부 Gateway를 분리하는 구성이 가장 깔끔합니다.
  • 버전별 정확한 Envoy 패치 버전은 projectcontour.io의 Compatibility Matrix에서 확인하세요 — 같은 Contour 마이너 버전이라도 패치마다 번들 Envoy가 보안 업데이트로 올라갑니다.

PROXY 헤더 없는 내부 호출 허용 & 대체 방법

여기서는 PROXY 헤더 없는 내부 호출 허용 & 대체 방법을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Proxy Protocol을 켠 뒤 가장 흔하게 터지는 문제는 LB를 거치지 않고 Envoy로 들어오는 호출입니다. 대표적으로 세 가지 경로가 있습니다.

① 클러스터 내부 Pod → Envoy Service(ClusterIP) 직접 호출 — 예: 사내 API 게이트웨이 정책을 타기 위해 envoy.projectcontour.svc.cluster.local로 요청. PROXY 헤더가 없으므로 Envoy가 연결을 끊습니다.
② 클러스터 내부 Pod → 외부 도메인(LB IP) 호출(헤어핀) — LB가 Service status.loadBalancer.ingress에 IP를 기록하면 kube-proxy가 그 IP로 가는 패킷을 노드에서 가로채 LB를 거치지 않고 Envoy로 바로 보냅니다. 결과적으로 PROXY 헤더 없이 도착해 실패합니다. AWS NLB처럼 hostname만 기록하는 LB에서는 보통 발생하지 않고, DigitalOcean·Hetzner·일부 온프레미스 LB처럼 IP를 기록하는 환경에서 발생합니다.
③ 사내망/VPN → 내부 LB → Envoy — 내부 LB에 Proxy Protocol을 켜지 않았다면 역시 실패합니다.

Envoy의 allow_requests_without_proxy_protocol
Envoy의 proxy_protocol 리스너 필터에는 allow_requests_without_proxy_protocol: true 옵션이 있어, 헤더가 있으면 읽고 없으면 그냥 통과시키는 "선택적 허용"이 가능합니다. Contour 1.28 ~ 1.33이 번들하는 Envoy는 모두 이 필드를 지원합니다. 하지만 Contour는 이 필드를 노출하지 않습니다 — Contour 설정 파일·contour serve 플래그·ContourConfiguration/ContourDeployment 어디에도 대응 옵션이 없고, Contour가 xDS로 리스너를 계속 재생성하므로 Istio의 EnvoyFilter 같은 임의 패치 경로도 없습니다. 따라서 Contour 1.33까지는 "PROXY 있으면 읽고 없어도 허용"을 설정으로 만들 수 없고, 아래 대체 방법 중 하나를 선택해야 합니다(새 버전에서 옵션이 추가되었는지는 Contour 릴리스 노트로 확인).

⚠️
allow 옵션을 쓸 수 있더라도 주의할 점
Envoy 문서는 이 옵션이 Proxy Protocol 스펙 준수를 깨뜨리므로 "리스너로 들어오는 모든 트래픽이 신뢰할 수 있는 출처일 때만" 켜라고 경고합니다. 헤더 없는 연결은 TCP 소스 IP(노드/SNAT IP)가 클라이언트 IP로 기록되므로 IP 기반 허용 정책이 섞일 수 있고, 반대로 외부에서 LB를 우회해 직접 들어온 연결이 PROXY 헤더를 위조하면 임의의 IP로 가장할 수 있습니다. 또한 v2 시그니처와 일치하는 12바이트 이하(v1은 6바이트 이하)의 아주 짧은 요청은 Envoy가 헤더인지 판단하지 못해 타임아웃됩니다. Envoy 리스너(NodePort/Pod IP)로의 직접 접근은 NetworkPolicy·보안그룹으로 LB와 신뢰 대역만 허용하세요.
다이어그램 렌더링 중…
option2-internal-gateway.yamlYAML
# [② 권장] 내부 전용 Gateway — Proxy Protocol OFF + ClusterIP Service
apiVersion: projectcontour.io/v1alpha1
kind: ContourDeployment
metadata:
  name: contour-internal
  namespace: projectcontour
spec:
  runtimeSettings:
    envoy:
      listener:
        useProxyProtocol: false          # 내부용은 헤더 없이 받음
  envoy:
    networkPublishing:
      type: ClusterIPService             # 클러스터 내부에서만 접근
      # 사내망/VPN에도 노출하려면 LoadBalancerService + internal LB annotation 사용:
      # type: LoadBalancerService
      # serviceAnnotations:
      #   service.beta.kubernetes.io/aws-load-balancer-scheme: internal
---
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
  name: contour-internal
spec:
  controllerName: projectcontour.io/gateway-controller
  parametersRef:
    group: projectcontour.io
    kind: ContourDeployment
    name: contour-internal
    namespace: projectcontour
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: internal-gateway
  namespace: projectcontour
spec:
  gatewayClassName: contour-internal
  listeners:
  - name: http
    protocol: HTTP
    port: 80
    allowedRoutes:
      namespaces:
        from: All
---
# 하나의 HTTPRoute를 외부(PP on) + 내부(PP off) Gateway 양쪽에 연결 → 라우팅 규칙 1벌로 관리
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: api-route
  namespace: production
spec:
  parentRefs:
  - name: external-gateway          # 이전 섹션의 PP on Gateway
    namespace: projectcontour
  - name: internal-gateway
    namespace: projectcontour
  hostnames:
  - "api.aidevops.kr"
  rules:
  - backendRefs:
    - name: api-service
      port: 8080

# 내부 Pod는 Provisioner가 만든 내부 envoy Service로 호출:
#   kubectl -n projectcontour get svc   # envoy-internal-gateway 형태의 이름 확인
#   curl -H "Host: api.aidevops.kr" http://envoy-internal-gateway.projectcontour.svc.cluster.local/
option2-static-second-instance.yamlYAML
# [② 정적 매니페스트 버전] Provisioner 없이 두 번째 Contour+Envoy 세트를 운영하는 경우
# - 두 번째 세트는 별도 네임스페이스(예: projectcontour-internal)에 quickstart 매니페스트를 복제 배포
# - contour serve 인자로 담당 IngressClass를 구분하고, PP 플래그는 외부 세트에만 추가
#     외부: contour serve ... --ingress-class-name=contour-external --use-proxy-protocol
#     내부: contour serve ... --ingress-class-name=contour-internal
apiVersion: projectcontour.io/v1
kind: HTTPProxy
metadata:
  name: api-internal
  namespace: production
spec:
  ingressClassName: contour-internal     # 내부 세트만 이 HTTPProxy를 처리
  virtualhost:
    fqdn: api.internal.aidevops.kr
  routes:
  - conditions:
    - prefix: /
    services:
    - name: api-service
      port: 8080
option3-client-sends-proxy-header.shBASH
# [③] 클라이언트가 직접 PROXY 헤더를 보내는 방법
# curl — 테스트·헬스체크 스크립트용 (v1 헤더 전송)
curl --haproxy-protocol -H "Host: api.aidevops.kr" \
  http://envoy.projectcontour.svc.cluster.local/healthz

# curl 8.2+ — 헤더에 넣을 클라이언트 IP 지정
curl --haproxy-protocol --haproxy-clientip 10.1.2.3 -H "Host: api.aidevops.kr" \
  http://envoy.projectcontour.svc.cluster.local/

# 애플리케이션 트래픽은 HTTP 라이브러리가 PROXY 헤더를 지원하지 않는 경우가 대부분이므로
# 같은 Pod에 TCP 프록시 사이드카를 두고 localhost로 호출하게 합니다 (HAProxy 예시):
#   frontend local_in
#       mode tcp
#       bind 127.0.0.1:18080
#       default_backend envoy
#   backend envoy
#       mode tcp
#       server contour envoy.projectcontour.svc.cluster.local:80 send-proxy-v2
# 앱은 http://127.0.0.1:18080 으로 요청 + Host 헤더 지정
option4-hairpin.yamlYAML
# [④] 헤어핀 차단 — 내부 Pod가 외부 도메인을 호출해도 반드시 LB를 거치게 만들기

# (a) Kubernetes 1.32+ (LoadBalancerIPMode GA): Cloud Controller가 ipMode: Proxy를 기록하면
#     kube-proxy가 LB IP를 노드에서 가로채지 않고 실제 LB로 보냄 → PROXY 헤더가 붙음
#     확인:
#     kubectl -n projectcontour get svc envoy -o jsonpath='{.status.loadBalancer.ingress}'
#     [{"ip":"203.0.113.10","ipMode":"Proxy"}]   ← Proxy면 OK, VIP 또는 미표시면 헤어핀 발생
#     ipMode는 사용자가 아니라 Cloud Controller Manager가 설정하는 값 — CCM 지원 여부 확인

# (b) IP 대신 hostname을 status에 기록하게 하는 LB별 annotation
apiVersion: v1
kind: Service
metadata:
  name: envoy
  namespace: projectcontour
  annotations:
    # DigitalOcean
    service.beta.kubernetes.io/do-loadbalancer-enable-proxy-protocol: "true"
    service.beta.kubernetes.io/do-loadbalancer-hostname: "lb.aidevops.kr"
    # Hetzner Cloud
    # load-balancer.hetzner.cloud/uses-proxyprotocol: "true"
    # load-balancer.hetzner.cloud/hostname: "lb.aidevops.kr"
spec:
  type: LoadBalancer
  selector:
    app: envoy
  ports:
  - name: https
    port: 443
    targetPort: 8443
option6-envoy-gateway.yamlYAML
# [⑥] Envoy Gateway로 이전 시 — Proxy Protocol 활성화 + 헤더 없는 연결 선택적 허용
# 1) ClientTrafficPolicy로 Proxy Protocol 활성화
#    (최신 Envoy Gateway는 proxyProtocol 블록의 optional 설정을 제공 — 사용 버전의 API 문서 확인)
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
  name: enable-proxy-protocol
  namespace: envoy-gateway-system
spec:
  targetRefs:
  - group: gateway.networking.k8s.io
    kind: Gateway
    name: eg
  enableProxyProtocol: true
---
# 2) optional 필드가 없는 버전이라면 EnvoyPatchPolicy로 Envoy 필드를 직접 패치
#    (EnvoyGateway 설정에 extensionApis.enableEnvoyPatchPolicy: true 필요)
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyPatchPolicy
metadata:
  name: pp-allow-without-header
  namespace: envoy-gateway-system
spec:
  targetRef:
    group: gateway.networking.k8s.io
    kind: Gateway
    name: eg
  type: JSONPatch
  jsonPatches:
  - type: "type.googleapis.com/envoy.config.listener.v3.Listener"
    name: envoy-gateway-system/eg/http        # <namespace>/<gateway>/<listener>
    operation:
      op: add
      path: "/listener_filters/0/typed_config/allow_requests_without_proxy_protocol"
      value: true
# 적용 후 egctl config envoy-proxy listener 로 listener_filters 순서(0번이 proxy_protocol인지) 확인
대체 방법해결하는 경로장점단점
① 내부 호출은 앱 Service 직접 호출 (svc.cluster.local)Pod → Envoy가장 단순, 추가 리소스 없음, 홉 1개 감소Envoy의 Rate limit·IP 정책·TLS·라우팅 규칙을 거치지 않음
② 내부 전용 Gateway(Envoy) 분리 — 권장Pod → Envoy, 사내망 → 내부 LBHTTPRoute 하나를 외부/내부 Gateway에 동시 연결해 라우팅 규칙 재사용Contour+Envoy 세트 추가 운영(리소스·모니터링)
③ 클라이언트가 PROXY 헤더를 직접 전송Pod → Envoy기존 Envoy 재사용일반 HTTP 클라이언트 라이브러리는 대부분 미지원 — 중간 TCP 프록시(HAProxy/Envoy 사이드카) 필요
④ 헤어핀 차단 (ipMode: Proxy / hostname annotation)Pod → 외부 도메인내부 트래픽도 LB를 거쳐 PROXY 헤더가 붙음Cloud Controller 지원 필요, LB 경유로 지연·비용 증가
⑤ Proxy Protocol 대신 externalTrafficPolicy: Local모든 경로 (PP 자체를 제거)내부 호출 문제가 원천적으로 사라짐LB가 패스스루(SNAT 없음)일 때만 가능
⑥ Envoy Gateway로 이전 + 선택적 PP모든 경로allow_requests_without_proxy_protocol을 직접 켤 수 있음컨트롤러 교체(HTTPProxy는 이전 불가, HTTPRoute는 재사용 가능)
⑦ Contour 포크/패치 빌드모든 경로Contour 유지리스너 필터 생성 코드를 패치해 직접 빌드 — 업그레이드마다 재적용, 운영 부담 큼 (비권장)

Tip

  • Contour를 유지한다면 "외부 Gateway(PP on) + 내부 Gateway(PP off)" 분리가 가장 안전하고 예측 가능한 구성입니다. HTTPRoute의 parentRefs에 두 Gateway를 모두 걸면 라우팅 규칙을 중복 관리할 필요도 없습니다.
  • 내부 전용 Envoy를 만들 때 Rate limit·인증 정책도 외부와 동일하게 적용되는지 확인하세요. HTTPProxy의 ipAllowFilterPolicy는 내부 Envoy에서는 Pod IP 대역 기준으로 판정됩니다.
  • PROXY 헤더를 위조하면 원본 IP를 속일 수 있습니다. Envoy Pod/NodePort는 NetworkPolicy와 보안그룹으로 LB·신뢰 대역에서만 접근 가능하게 막아두세요.

Proxy Protocol 검증 & 트러블슈팅

여기서는 Proxy Protocol 검증 & 트러블슈팅을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Proxy Protocol 문제는 대부분 "LB와 Envoy의 설정 불일치" 또는 "LB를 거치지 않는 경로" 둘 중 하나입니다. 증상만 보면 네트워크 장애처럼 보이지만, 증상의 형태로 어느 쪽이 켜져 있는지 거의 정확히 추론할 수 있습니다. 아래 표로 원인을 좁힌 뒤, 코드 블록의 순서대로 Envoy 설정 → 헤더 유무별 요청 → 통계 → 액세스 로그를 확인하세요.
pp-verify.shBASH
NS=projectcontour
ENVOY_POD=$(kubectl -n $NS get pod -l app=envoy -o jsonpath='{.items[0].metadata.name}')

# 1) Envoy 리스너에 proxy_protocol 필터가 붙었는지
kubectl -n $NS port-forward "$ENVOY_POD" 9001:9001 >/dev/null &
PF_PID=$!; sleep 2
curl -s localhost:9001/config_dump \
  | grep -B2 -A4 "envoy.filters.listener.proxy_protocol"

# 2) Proxy Protocol 관련 통계 — 헤더 파싱 실패/거부 카운터가 증가하는지
curl -s localhost:9001/stats | grep -i proxy_proto
kill $PF_PID

# 3) 헤더 유무별 요청 비교 (클러스터 내부 디버그 Pod)
kubectl run pp-test -n default --rm -it --restart=Never --image=curlimages/curl -- sh -c '
  echo "--- PROXY 헤더 없이 (PP on이면 실패가 정상)";
  curl -sv -m 5 -H "Host: api.aidevops.kr" http://envoy.projectcontour.svc.cluster.local/ 2>&1 | tail -3;
  echo "--- PROXY v1 헤더 포함 (PP on이면 성공해야 정상)";
  curl -sv -m 5 --haproxy-protocol -H "Host: api.aidevops.kr" http://envoy.projectcontour.svc.cluster.local/ 2>&1 | tail -3'

# 4) 외부에서 요청 후 Envoy 액세스 로그에 실제 공인 IP가 찍히는지
curl -s https://api.aidevops.kr/ -o /dev/null
kubectl -n $NS logs "$ENVOY_POD" -c envoy --tail=5
# 기본 로그 포맷의 x-forwarded-for / downstream 주소가 내 공인 IP여야 정상

# 5) 헤어핀 여부 — status에 IP가 있고 ipMode가 Proxy가 아니면 내부 호출은 LB를 우회
kubectl -n $NS get svc -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.loadBalancer.ingress}{"\n"}{end}'
envoy-networkpolicy.yamlYAML
# PROXY 헤더 위조 방지 — Envoy 트래픽 포트는 LB가 있는 VPC 대역에서만 허용
# (ip 타깃 NLB는 VPC 서브넷 IP에서 연결, 내부 호출은 내부 Gateway로 분리했다는 전제)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: envoy-allow-lb-only
  namespace: projectcontour
spec:
  podSelector:
    matchLabels:
      app: envoy
  policyTypes: ["Ingress"]
  ingress:
  - from:
    - ipBlock:
        cidr: 10.0.0.0/16          # LB 서브넷 / VPC CIDR
    ports:
    - port: 8080
    - port: 8443
  - from:                          # Prometheus 메트릭 수집(8002)
    - namespaceSelector:
        matchLabels:
          kubernetes.io/metadata.name: monitoring
    ports:
    - port: 8002
증상원인확인 / 해결
외부 요청 전부 실패 — curl "Empty reply from server", "Connection reset", HTTPS는 핸드셰이크 도중 끊김Envoy만 PP on, LB는 offLB annotation / 타깃 그룹 proxy_protocol_v2.enabled 확인
HTTP는 400 Bad Request, HTTPS는 "wrong version number"·SSL 오류LB만 PP on, Envoy는 off (PROXY 헤더가 HTTP/TLS 데이터로 해석됨)config_dump에 proxy_protocol 필터 존재 여부 확인 → useProxyProtocol 적용
외부는 정상, 내부 Pod → Envoy Service 호출만 실패PROXY 헤더 없는 내부 호출내부 전용 Gateway 분리 또는 앱 Service 직접 호출
외부는 정상, 내부 Pod → 외부 도메인 호출만 실패kube-proxy 헤어핀 (LB IP를 노드에서 가로챔)Service status의 ipMode 확인, hostname annotation 적용
LB 타깃이 모두 unhealthy헬스체크 연결의 PROXY 헤더 유무와 대상 포트 불일치트래픽 포트 TCP 헬스체크로 전환 후 재확인
정상 동작하지만 로그의 클라이언트 IP가 여전히 노드/LB IPPP 미적용, 또는 앞단에 L7 프록시가 한 단계 더 있음config_dump 확인, L7이 있다면 numTrustedHops를 홉 수만큼 설정
아주 짧은 요청이 간헐적으로 타임아웃 (선택적 허용 사용 시)12바이트(v1: 6바이트) 이하 요청을 헤더 대기 중으로 판단allow 옵션 대신 리스너 분리 구성으로 전환

Tip

  • 전환 작업 체크리스트: ① LB가 SNAT을 하는지 확인 → ② 헤더 없는 내부 호출 경로 목록화 → ③ 내부 Gateway 분리 또는 대체 방법 결정 → ④ 스테이징에서 LB·Envoy 동시 적용 → ⑤ config_dump·액세스 로그로 원본 IP 확인 → ⑥ 운영은 새 LB + DNS 전환으로 무중단 적용.
  • Contour 업그레이드(특히 1.33.5 → 1.33.6처럼 번들 Envoy 마이너 버전이 바뀌는 패치) 후에는 pp-verify.sh를 다시 돌려 리스너 필터와 원본 IP 기록이 유지되는지 확인하세요.

Contour/Envoy 모니터링 — Prometheus & Grafana

여기서는 Contour/Envoy 모니터링 — Prometheus & Grafana을 실제 코드와 함께 확인합니다. 예제를 그대로 따라 하기보다, 입력과 출력, 그리고 바뀌기 쉬운 부분이 어디인지 보면서 읽어보세요.

Contour는 자체 컨트롤 플레인 메트릭을 :8000/metrics에 노출하고, 데이터 플레인인 Envoy는 :8002/stats/prometheus에서 xDS 반영 상태·업스트림 커넥션·요청 처리 결과를 노출합니다. kube-prometheus-stack을 이미 설치했다면(모니터링 스택 구성 참고) ServiceMonitor 2개만 추가하면 됩니다. Envoy의 admin 인터페이스(9001)는 클러스터 내부에서만 접근하도록 외부 노출을 반드시 차단하세요.
servicemonitor-contour-envoy.yamlYAML
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: contour
  namespace: monitoring
  labels:
    release: kube-prometheus-stack   # Helm 릴리스 이름과 일치해야 수집됨
spec:
  namespaceSelector:
    matchNames: ["projectcontour"]
  selector:
    matchLabels:
      app: contour
  endpoints:
  - port: metrics
    path: /metrics
    interval: 30s
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: envoy
  namespace: monitoring
  labels:
    release: kube-prometheus-stack
spec:
  namespaceSelector:
    matchNames: ["projectcontour"]
  selector:
    matchLabels:
      app: envoy
  endpoints:
  - port: metrics
    path: /stats/prometheus
    interval: 30s
prometheusrule-envoy.yamlYAML
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: contour-envoy-alerts
  namespace: monitoring
  labels:
    release: kube-prometheus-stack
spec:
  groups:
  - name: contour.rules
    rules:
    - alert: EnvoyHigh5xxRate
      expr: sum(rate(envoy_cluster_upstream_rq_xx{envoy_response_code_class="5"}[5m])) / sum(rate(envoy_cluster_upstream_rq_xx[5m])) > 0.05
      for: 5m
      labels: { severity: critical }
      annotations:
        summary: "Envoy 5xx 응답 비율 5% 초과"

    - alert: EnvoyUpstreamConnectFail
      expr: increase(envoy_cluster_upstream_cx_connect_fail[5m]) > 20
      for: 5m
      labels: { severity: warning }
      annotations:
        summary: "Envoy 업스트림 커넥션 실패 급증 — 백엔드 Pod 헬스체크 확인 필요"

    - alert: ContourReconcileErrors
      expr: increase(contour_dagrebuild_total{result="error"}[15m]) > 0
      for: 5m
      labels: { severity: warning }
      annotations:
        summary: "Contour DAG 재구성 오류 발생 — HTTPProxy/HTTPRoute 설정 충돌 여부 확인"

    - alert: WildcardCertExpiringSoon
      expr: (certmanager_certificate_expiration_timestamp_seconds{name="wildcard-aidevops-kr"} - time()) / 86400 < 14
      for: 1h
      labels: { severity: warning }
      annotations:
        summary: "와일드카드 인증서 만료 14일 이내 — DNS-01 갱신 실패 여부 확인"
grafana-dashboards.shBASH
# Envoy 공식 Grafana 대시보드 (grafana.com/dashboards) — Import ID 입력만으로 사용 가능
# Envoy Global:   Dashboard ID 11021
# Envoy Clusters: Dashboard ID 11022

# Contour 자체 대시보드는 grafana.com에 등록되어 있지 않으므로 저장소에서 JSON을 직접 가져옵니다
git clone --depth 1 https://github.com/projectcontour/contour.git /tmp/contour
# /tmp/contour/examples/grafana/ 하위 JSON을 Grafana UI → Dashboards → Import 로 업로드
# (버전에 따라 경로가 변경될 수 있으니 실제 clone한 태그의 examples/ 디렉터리를 확인)

다음 단계

이 섹션은 다음 단계을 실무 관점에서 정리합니다. 개념을 외우기보다, 어떤 상황에서 이 기준을 꺼내 쓸지에 초점을 맞춰보세요.

🚀
K8s 심화 이후 로드맵

• GitOps: ArgoCD / Flux로 Git 기반 선언적 배포 자동화
• 서비스 메시: Istio / Linkerd — 트래픽 제어, mTLS, 관찰 가능성
• CKA / CKAD 자격증: 실무 K8s 역량 공인 자격증

연계 가이드: Kubernetes 기초 가이드 · CI/CD 가이드 · Prometheus 가이드 · Grafana 가이드
← 이전 가이드Kubernetes다음 가이드 →Prometheus