Lộ trình
Cloud Native Computing FoundationAssociateTuần 1: Nền tảngBài 2 / 30

Ngày 2: Kiến trúc Kubernetes

Thời lượng: 45 phút
Mục tiêu: 2 nhiệm vụ chính
Tiến độ lộ trình
ckangày 2
hoàn thành2 / 30 bài
Bối cảnh bài học

Đọc kiến trúc Kubernetes qua request lifecycle: API server, etcd, scheduler, controllers, kubelet, runtime và kube-proxy.

đọc hiểuthực hànhcheckpoint
Bài giảng hôm nay

Học hiểu, rồi mới thực hành

Đọc kiến trúc Kubernetes qua request lifecycle: API server, etcd, scheduler, controllers, kubelet, runtime và kube-proxy.

Bắt đầu đọc bài giảng

Nhiệm vụ bài học hôm nay

  • Học vai trò control plane: API Server, Scheduler, Controller Manager, etcd
  • Học vai trò node: kubelet, kube-proxy, container runtime
Instructor walkthrough

Bài giảng chi tiết: từ bài toán đến bằng chứng

Scenario xuyên suốt

Một Pod ở trạng thái Pending, một node NotReady và một object apply thành công nhưng không được reconcile. Người học cần biết component nào chịu trách nhiệm, evidence nào cần thu và không chẩn đoán bằng cách restart tùy ý.

01Đọc bài toán

Xác định actor, workload, constraint và trạng thái cuối cần đạt.

02Vẽ luồng / boundary

Chỉ ra request, dependency, identity và failure domain trước khi chọn công cụ.

03Chọn và thực hành

Thay đổi nhỏ nhất trong lab cô lập; command nào cũng phải nói rõ nó kiểm tra điều gì.

04Kiểm chứng / recovery

Đối chiếu trạng thái thực tế, tạo một failure variant và ghi cách hoàn tác.

Cách nối lý thuyết với thực tế
  • API server là front door/authentication/validation/watch hub; etcd lưu cluster state; không coi etcd là application database.
  • Scheduler chọn node cho Pod chưa bind; controller manager reconcile desired/actual state; kubelet thực thi Pod trên node.
  • Container runtime chạy container; kube-proxy/service datapath xử lý network behavior tùy implementation; component roles không đồng nhất.
  • Control-plane failure, node failure, scheduling failure và application failure có symptom/evidence khác nhau.

Từ desired state đến running Pod

kubectl gửi request tới API server. API server validate/auth rồi ghi state vào etcd và phát watch; scheduler chọn node cho Pod chưa bind; controller đảm bảo resource phụ; kubelet trên node gọi runtime, cập nhật status; service datapath xử lý network. Khi debug, hỏi state dừng ở bước nào.

Failure matrix

Pending có thể do scheduler/resource/taint; NotReady do kubelet/runtime/network/condition; ImagePullBackOff do registry/image/secret; API timeout do control-plane/network/auth. Thu evidence bằng get/describe/events/logs trước khi sửa.

Bài tập

Vẽ lifecycle và lập failure matrix năm tình huống. Dùng cluster lab để đọc events/node conditions, ghi expected state và cleanup. Pass khi giải thích được “component nào biết gì” và chọn được command đầu tiên có giá trị nhất.

Terminal reference

Command list và cách dùng

Chạy từng lệnh theo đúng thứ tự. Trước các lệnh có thể tạo hoặc thay đổi tài nguyên, hãy kiểm tra profile, account và region.

Commands · read-only checkpoints
kubectl get pods -A -o wide
kubectl get nodes -o wide && kubectl describe node NODE_NAME
kubectl get events -A --sort-by=.lastTimestamp
kubectl get componentstatuses
kubectl version --output=yaml
Hands-on lab

Thực hành theo scenario

  1. Vẽ flow `kubectl apply → API server → etcd/watch → scheduler/controller → kubelet/runtime → status`; ghi component và expected state từng bước.
  2. Dùng `kubectl get/describe/events`, node conditions, component/version signals để phân loại Pending, NotReady và reconcile delay.
  3. Tạo design-only failure matrix: API unreachable, scheduler constraint, image pull, kubelet/runtime, CNI/service; với mỗi lỗi ghi symptom/evidence/remediation/rollback.
  4. Không kill control-plane process hoặc sửa etcd production. Cleanup chỉ xóa marker/workload lab theo context/namespace đã xác nhận.
Evidence checkpoint

Kiểm chứng kết quả

Không coi lệnh chạy thành công là đủ. Hãy đối chiếu output với trạng thái mong đợi:

  • Component map đúng và nối được lifecycle.
  • Phân biệt control-plane/node/scheduling/application failure.
  • Có evidence command và expected state cho từng symptom.
  • Troubleshooting ưu tiên read-only và có rollback.
  • Cleanup lab scoped, không đụng cluster production.
Transfer to exam / production

Bẫy thường gặp và trade-off

CKA thường yêu cầu tìm nhanh component/symptom: Pod Pending không đồng nghĩa app lỗi; Node NotReady không sửa bằng scale; API server/etcd/scheduler có evidence khác nhau.

Checkpoint · 3 phút

Kiểm tra nhanh

Câu hỏi: Pod ở Pending nhưng API server trả lời bình thường. Kiểm tra đầu tiên phù hợp nhất là gì?

Kết thúc bài

Checklist trước khi sang Ngày 2