LODY/정리

클라우드 네이티브 / 관리형 Kubernetes 배포를 CKA 기준으로 다시 읽기

관리형 Kubernetes 배포를 CKA 기준으로 다시 읽기

Control Plane과 Data Plane

Kubernetes Cluster는 Control PlaneData Plane으로 나뉜다.

  • Control Plane: API Server, etcd, Scheduler, Controller Manager. 원하는 상태를 받아 Pod를 schedule하고 Controller loop로 맞춘다
  • Data Plane: Worker Node 위의 kubelet, kube-proxy, Container Runtime, Pod. Workload가 실제로 도는 쪽이다
API ServeretcdSchedulerController Mgretcd Backup·QuorumTLS CertificateVersion UpgradekubeletPods / CSIscheduleControl PlaneMaster Node운영 부담Data PlaneWorker Node

Control Plane 컴포넌트는 보통 Master Node에서 돈다. Cluster를 직접 올리면 Master Node를 사용자가 준비하고, 아래 운영도 같이 한다.

etcd

Cluster 상태 저장소다. Backup, 복구, Quorum을 직접 관리한다.

TLS Certificate

API Server, kubelet, etcd 사이를 TLS로 묶는다. Certificate가 만료되거나 Rotation을 놓치면 Cluster가 멈춘다.

Upgrade

Control Plane 버전을 올릴 때는 컴포넌트 순서와 호환을 맞춰야 한다.

위 운영까지 포함한 구성이 Self-hosted Control Plane이다.


관리형 Kubernetes

관리형에서는 Control Plane 운영을 클라우드가 맡는다. 사용자는 Data Plane만 보면 된다.

API ServeretcdSchedulerController Mgretcd Backup·QuorumTLS CertificateVersion UpgradekubeletPods / CSIscheduleControl Plane (클라우드가 운영)Master Node운영 (클라우드)Data Plane (내가 운영)Worker Node

클라우드마다 상품 이름은 다르다.

  • AWS: EKS
  • Azure: AKS
  • NHN Cloud: NKS (NHN Kubernetes Service)

다룰 내용:

  • Worker Node
  • Security Group
  • Workload Manifest
  • StorageClass
  • image pull authentication
  • external LB·DNS

실습 구성

NHN Cloud NKS에 API·Worker·Web·Postgres·Caddy를 한 Cluster에 올린다.

  • DB까지 Cluster 안, 1 Node
  • Postgres는 CloudNativePG(CNPG). 나중에 Worker를 늘린 뒤 instances를 올린다
  • API는 replicas=1, 부하는 Worker로 뺀다

Kubernetes 리소스:

구분기술역할
클라우드NHN Cloud 문서VPC, Load Balancer, Worker Node Security Group
Orchestration관리형 Kubernetes (NHN NKS)EKS·AKS와 같은 계열. Workload·Secret·Service 실행 환경. 사용 가이드
Container RegistryNCRamd64 이미지 보관·pull. 사용 가이드
CIGitHub ActionsNative amd64 빌드 후 NCR push
EdgeCaddy (Automatic HTTPS)단일 도메인 HTTPS, path routing, ACME
DNSDNS A Record (Proxy off)도메인에서 LB로. TLS는 Edge에서 종료
DBCloudNativePG (Postgres)Cluster 안 Postgres. 당시 instances=1
스키마Job (migrate)빈 DB에 upgrade로 스키마 적용
StorageCSI (Cinder CSI) + StorageClass (fstype: ext4)CNPG PVC. 블록은 NHN Block Storage. NHN에선 SC를 Manifest로 직접 정의
Backuppg_dumpObject Storage1 Node 전제의 최소 Backup
InternetDNS (A only)NHN LBCaddyHTTPSAPI ×1WebWorkermigratePostgresCNPG ×1HTTPS/api/v1/*그 외schemaNHN CloudVPC / public subnetmanaged k8s clusterworkloadsdata

Workload: Deployment와 Job

Pod는 schedule되는 최소 단위다. 운영에서는 Pod를 직접 오래 두지 않고 Controller에 맡긴다.

  • Deployment: ReplicaSet으로 replica 수와 이미지를 선언하면 Controller가 유지한다
  • Job: 완료가 목표다. 스키마 migration처럼 한 번 성공하면 끝난다

프로세스에 Singleton 전제가 있으면 replicas만 올리면 안 된다.

Workload종류역할
PostgresCNPG Cluster (당시 1 Instance)앱 DB. Primary Service(*-rw)로 접속
migrateJob (일회성)빈 DB에 스키마 적용
APIDeployment (replicas=1)HTTP + Coordinator 작업. Singleton
WorkerDeployment (1+)백그라운드. 수평 증설 가능
WebDeploymentFrontend 정적/SSR
CaddyDeployment + LoadBalancer단일 도메인 HTTPS Edge

API에 Coordinator 작업이 남아 있어서 replicas=1이다. 부하는 Worker로 뺀다. CNPG도 당시 instances=1이다.


Service와 Networking

Service는 Pod IP가 바뀌어도 같은 이름으로 붙는 Endpoint다.

  • ClusterIP: Cluster 안에서만 도달
  • NodePort: 각 Node의 고정 포트로 노출. 클라우드 LB health check가 이 경로를 쓰는 경우가 많다
  • LoadBalancer: 클라우드가 외부 LB를 만들고, 보통 뒤에서 NodePort로 Node에 붙는다

Ingress + cert-manager 대신 LB 뒤 Caddy에서 TLS를 끝내고, Certificate는 Caddy ACME로 받는다. 단일 도메인에서 API와 웹을 path로 나눈다. API는 CNPG Primary(*-rw)에만 붙인다. DNS Proxy가 끼면 ACME HTTP-01/TLS-ALPN이 깨질 수 있다.

ClientDNS ALB :443CaddyAPI :8080Webpg-rw/api/v1/*그 외clusteredgeapps

DNS는 A Record만 쓰고 Proxy는 끈다. Caddy가 origin에서 ACME를 처리할 때 앞단 Proxy가 challenge를 가로채면 발급이 실패할 수 있다.

NHN LB는 NodePort로 health check한다. Worker Security Group이 그 대역을 막으면 멤버가 DOWN이다.

NHN LBNode SGTCP 30000-32767Worker NodePortSG 미개방 시 멤버 DOWNhealth checkVPC / worker node path
  • 공인 LB + Caddy HTTPS + DNS only
  • Worker Node Security Group에 LB health check용 NodePort 대역(TCP 30000-32767) inbound 개방

Service type만으로는 부족하다. 클라우드 LB와 Node Security Group도 Data Plane이다.


Storage: PVC와 권한

  • PersistentVolumeClaim (PVC): Workload가 요청하는 Volume
  • PersistentVolume (PV): 실제로 붙는 Volume
  • StorageClass (SC): CSI 드라이버와 parameter로 동적 provisioning

이 Cluster의 cinder-csi는 fsGroupPolicyReadWriteOnceWithFSType이다. StorageClass에 fstype이 있어야 fsGroup이 적용되고, 없으면 Postgres uid 쓰기가 Permission denied다. Volume Bound와 프로세스 쓰기 가능은 다르다.

NHN NKS + cinder-csi에서는 StorageClass가 자동 생성되지 않았다. SC를 Manifest로 정의하고 fstype: ext4를 명시한다. CNPG Operator는 Pod·Service·PVC를 만들지만 SC는 만들지 않는다.

빈 DB 스키마는 migrate Job이 먼저 upgrade로 만들고, API는 그다음 기동한다. API Entrypoint가 Stamp 우선이면 빈 DB에서 테이블 없이 Stamp만 찍힐 수 있다.

StorageClasscinder + ext4PVCCNPG Clustermigrate Jobalembic upgradeAPI bootreadyschema okdata planeblock storagebootstrap order

SC·fstype이 빠지면 PVC Pending이나 EPERM이다. 리비전 ID가 기본 alembic_version.version_num보다 길면 insert가 실패하고 같은 오류가 반복된다. 빈 DB에서는 버전 테이블을 넉넉히 미리 만들고 Job을 돌린다.

배포 순서:

  1. StorageClass
  2. CNPG Operator, Secret, NCR pull secret
  3. CNPG Cluster, migrate Job
  4. API, Worker, Web
  5. DNS 전파 확인 후 Caddy HTTPS

Security: pull authentication과 Secret

  • Secret: 민감 값 오브젝트. at-rest 암호화는 Cluster 설정에 달리고, 깃에 넣지 않는다
  • ServiceAccount: Pod가 API·Registry에 쓰는 신원
  • imagePullSecrets: private Registry면 SA 또는 Pod에 붙여야 kubelet이 이미지를 받는다

같은 프로젝트 Registry라도 pull authentication은 자동이 아니다.

  • runtime Secret은 Cluster Secret으로만 주입한다
  • NCR pull secret을 만들어 default ServiceAccount에 붙인다
  • NCR Console “미인증 이미지 Pull 방지”가 켜져 있으면 서명 없는 이미지가 412로 거절된다

앱 Lifecycle: 이미지와 재배포

  • Node Architecture와 Image Architecture가 다르면 schedule은 돼도 Container가 안 뜬다
  • Registry 정책(서명, provenance, Cache Manifest)은 kubelet pull 실패로만 보인다
  • 재배포 최소 단위는 이미지 Tag다

Cluster는 amd64 Node다. GitHub Actions에서 Native amd64로 빌드해 NCR에 push한다.

  • buildx Cache·provenance 첨부 Manifest를 NCR이 거부할 수 있다. provenance는 끈다
  • “미인증 이미지 Pull 방지”가 켜져 있으면 서명 없는 이미지가 412다
  • pull secret을 default ServiceAccount에 붙인다
git / workflowGitHub Actionsamd64 buildNCRSA + pull secretAPI / Web / Workerpushpullbuild & registrycluster

Manifest는 Registry·Tag를 채워 kubectl apply한다. Tag 하나로 API·웹·Worker를 같이 갱신한다.

Frontend는 Node 22 + pnpm 메이저 고정이다. corepack 최신 pnpm과 pnpm 10 native 스크립트 차단이 sharp를 깨뜨릴 수 있다.

재배포는 TAG를 바꿔 Manifest를 다시 적용한다.


한계

  • 1 Node · CNPG instances=1이면 Node는 단일 장애점이다
  • Snapshot·Dump 없이 Node만 믿으면 안 된다
  • API Singleton이라 수평 확장은 Worker 쪽만 열린다
  • CNPG·Volume·Backup은 직접 운영한다

pg_dump 후 Object Storage에 둔다. Node를 늘린 뒤 CNPG instances / 웹 replicas 경로는 Manifest 주석으로 남겼다.

node-1pg primarynode-2node-3pg replicapg_dumpObject Storagebackuplater HA현재이후backup

정리

Manifest만으로 끝나지 않는 지점:

  • Service type
  • StorageClass
  • pull secret
  • Security Group

StorageClass·Registry·Security Group 기본값은 클라우드마다 다르니 배포 순서에 넣어 확인한다.