Control Plane과 Data Plane
Kubernetes Cluster는 Control Plane과 Data Plane으로 나뉜다.
- Control Plane: API Server, etcd, Scheduler, Controller Manager. 원하는 상태를 받아 Pod를 schedule하고 Controller loop로 맞춘다
- Data Plane: Worker Node 위의 kubelet, kube-proxy, Container Runtime, Pod. Workload가 실제로 도는 쪽이다
Control Plane 컴포넌트는 보통 Master Node에서 돈다. Cluster를 직접 올리면 Master Node를 사용자가 준비하고, 아래 운영도 같이 한다.
etcd
Cluster 상태 저장소다. Backup, 복구, Quorum을 직접 관리한다.
TLS Certificate
API Server, kubelet, etcd 사이를 TLS로 묶는다. Certificate가 만료되거나 Rotation을 놓치면 Cluster가 멈춘다.
Upgrade
Control Plane 버전을 올릴 때는 컴포넌트 순서와 호환을 맞춰야 한다.
위 운영까지 포함한 구성이 Self-hosted Control Plane이다.
관리형 Kubernetes
관리형에서는 Control Plane 운영을 클라우드가 맡는다. 사용자는 Data Plane만 보면 된다.
클라우드마다 상품 이름은 다르다.
- AWS: EKS
- Azure: AKS
- NHN Cloud: NKS (NHN Kubernetes Service)
다룰 내용:
- Worker Node
- Security Group
- Workload Manifest
- StorageClass
- image pull authentication
- external LB·DNS
실습 구성
NHN Cloud NKS에 API·Worker·Web·Postgres·Caddy를 한 Cluster에 올린다.
- DB까지 Cluster 안, 1 Node
- Postgres는 CloudNativePG(CNPG). 나중에 Worker를 늘린 뒤
instances를 올린다 - API는 replicas=1, 부하는 Worker로 뺀다
Kubernetes 리소스:
- Deployment: API, Worker, Web, Caddy
- Job: migrate (일회성)
- Service: Edge는 LoadBalancer
- Secret, ServiceAccount: runtime Secret, NCR pull
- StorageClass, PersistentVolumeClaim
- CloudNativePG Cluster CR
| 구분 | 기술 | 역할 |
|---|---|---|
| 클라우드 | NHN Cloud 문서 | VPC, Load Balancer, Worker Node Security Group |
| Orchestration | 관리형 Kubernetes (NHN NKS) | EKS·AKS와 같은 계열. Workload·Secret·Service 실행 환경. 사용 가이드 |
| Container Registry | NCR | amd64 이미지 보관·pull. 사용 가이드 |
| CI | GitHub Actions | Native amd64 빌드 후 NCR push |
| Edge | Caddy (Automatic HTTPS) | 단일 도메인 HTTPS, path routing, ACME |
| DNS | DNS A Record (Proxy off) | 도메인에서 LB로. TLS는 Edge에서 종료 |
| DB | CloudNativePG (Postgres) | Cluster 안 Postgres. 당시 instances=1 |
| 스키마 | Job (migrate) | 빈 DB에 upgrade로 스키마 적용 |
| Storage | CSI (Cinder CSI) + StorageClass (fstype: ext4) | CNPG PVC. 블록은 NHN Block Storage. NHN에선 SC를 Manifest로 직접 정의 |
| Backup | pg_dump 후 Object Storage | 1 Node 전제의 최소 Backup |
Workload: Deployment와 Job
Pod는 schedule되는 최소 단위다. 운영에서는 Pod를 직접 오래 두지 않고 Controller에 맡긴다.
- Deployment: ReplicaSet으로 replica 수와 이미지를 선언하면 Controller가 유지한다
- Job: 완료가 목표다. 스키마 migration처럼 한 번 성공하면 끝난다
프로세스에 Singleton 전제가 있으면 replicas만 올리면 안 된다.
| Workload | 종류 | 역할 |
|---|---|---|
| Postgres | CNPG Cluster (당시 1 Instance) | 앱 DB. Primary Service(*-rw)로 접속 |
| migrate | Job (일회성) | 빈 DB에 스키마 적용 |
| API | Deployment (replicas=1) | HTTP + Coordinator 작업. Singleton |
| Worker | Deployment (1+) | 백그라운드. 수평 증설 가능 |
| Web | Deployment | Frontend 정적/SSR |
| Caddy | Deployment + LoadBalancer | 단일 도메인 HTTPS Edge |
API에 Coordinator 작업이 남아 있어서 replicas=1이다. 부하는 Worker로 뺀다. CNPG도 당시 instances=1이다.
Service와 Networking
Service는 Pod IP가 바뀌어도 같은 이름으로 붙는 Endpoint다.
- ClusterIP: Cluster 안에서만 도달
- NodePort: 각 Node의 고정 포트로 노출. 클라우드 LB health check가 이 경로를 쓰는 경우가 많다
- LoadBalancer: 클라우드가 외부 LB를 만들고, 보통 뒤에서 NodePort로 Node에 붙는다
Ingress + cert-manager 대신 LB 뒤 Caddy에서 TLS를 끝내고, Certificate는 Caddy ACME로 받는다. 단일 도메인에서 API와 웹을 path로 나눈다. API는 CNPG Primary(*-rw)에만 붙인다. DNS Proxy가 끼면 ACME HTTP-01/TLS-ALPN이 깨질 수 있다.
DNS는 A Record만 쓰고 Proxy는 끈다. Caddy가 origin에서 ACME를 처리할 때 앞단 Proxy가 challenge를 가로채면 발급이 실패할 수 있다.
NHN LB는 NodePort로 health check한다. Worker Security Group이 그 대역을 막으면 멤버가 DOWN이다.
- 공인 LB + Caddy HTTPS + DNS only
- Worker Node Security Group에 LB health check용 NodePort 대역(TCP 30000-32767) inbound 개방
Service type만으로는 부족하다. 클라우드 LB와 Node Security Group도 Data Plane이다.
Storage: PVC와 권한
- PersistentVolumeClaim (PVC): Workload가 요청하는 Volume
- PersistentVolume (PV): 실제로 붙는 Volume
- StorageClass (SC): CSI 드라이버와 parameter로 동적 provisioning
이 Cluster의 cinder-csi는 fsGroupPolicy가 ReadWriteOnceWithFSType이다. StorageClass에 fstype이 있어야 fsGroup이 적용되고, 없으면 Postgres uid 쓰기가 Permission denied다. Volume Bound와 프로세스 쓰기 가능은 다르다.
NHN NKS + cinder-csi에서는 StorageClass가 자동 생성되지 않았다. SC를 Manifest로 정의하고 fstype: ext4를 명시한다. CNPG Operator는 Pod·Service·PVC를 만들지만 SC는 만들지 않는다.
빈 DB 스키마는 migrate Job이 먼저 upgrade로 만들고, API는 그다음 기동한다. API Entrypoint가 Stamp 우선이면 빈 DB에서 테이블 없이 Stamp만 찍힐 수 있다.
SC·fstype이 빠지면 PVC Pending이나 EPERM이다. 리비전 ID가 기본 alembic_version.version_num보다 길면 insert가 실패하고 같은 오류가 반복된다. 빈 DB에서는 버전 테이블을 넉넉히 미리 만들고 Job을 돌린다.
배포 순서:
- StorageClass
- CNPG Operator, Secret, NCR pull secret
- CNPG Cluster, migrate Job
- API, Worker, Web
- DNS 전파 확인 후 Caddy HTTPS
Security: pull authentication과 Secret
- Secret: 민감 값 오브젝트. at-rest 암호화는 Cluster 설정에 달리고, 깃에 넣지 않는다
- ServiceAccount: Pod가 API·Registry에 쓰는 신원
- imagePullSecrets: private Registry면 SA 또는 Pod에 붙여야 kubelet이 이미지를 받는다
같은 프로젝트 Registry라도 pull authentication은 자동이 아니다.
- runtime Secret은 Cluster Secret으로만 주입한다
- NCR pull secret을 만들어 default ServiceAccount에 붙인다
- NCR Console “미인증 이미지 Pull 방지”가 켜져 있으면 서명 없는 이미지가 412로 거절된다
앱 Lifecycle: 이미지와 재배포
- Node Architecture와 Image Architecture가 다르면 schedule은 돼도 Container가 안 뜬다
- Registry 정책(서명, provenance, Cache Manifest)은 kubelet pull 실패로만 보인다
- 재배포 최소 단위는 이미지 Tag다
Cluster는 amd64 Node다. GitHub Actions에서 Native amd64로 빌드해 NCR에 push한다.
- buildx Cache·provenance 첨부 Manifest를 NCR이 거부할 수 있다. provenance는 끈다
- “미인증 이미지 Pull 방지”가 켜져 있으면 서명 없는 이미지가 412다
- pull secret을 default ServiceAccount에 붙인다
Manifest는 Registry·Tag를 채워 kubectl apply한다. Tag 하나로 API·웹·Worker를 같이 갱신한다.
Frontend는 Node 22 + pnpm 메이저 고정이다. corepack 최신 pnpm과 pnpm 10 native 스크립트 차단이 sharp를 깨뜨릴 수 있다.
재배포는 TAG를 바꿔 Manifest를 다시 적용한다.
한계
- 1 Node · CNPG
instances=1이면 Node는 단일 장애점이다 - Snapshot·Dump 없이 Node만 믿으면 안 된다
- API Singleton이라 수평 확장은 Worker 쪽만 열린다
- CNPG·Volume·Backup은 직접 운영한다
pg_dump 후 Object Storage에 둔다. Node를 늘린 뒤 CNPG instances / 웹 replicas 경로는 Manifest 주석으로 남겼다.
정리
Manifest만으로 끝나지 않는 지점:
- Service type
- StorageClass
- pull secret
- Security Group
StorageClass·Registry·Security Group 기본값은 클라우드마다 다르니 배포 순서에 넣어 확인한다.