SRE Prep
เมนูโมดูล

9. Security & Testing Basics

ทำไมหัวข้อนี้สำคัญกับ interview

SRE ไม่ใช่ security engineer แต่ต้องดูแล production ให้ ปลอดภัยและถูกทดสอบ — เพราะ security misconfig และการ deploy ที่ไม่ทดสอบก็ทำให้เกิด outage ได้พอ ๆ กับ bug. คำถาม ระดับพื้นฐาน: "จัดการ secrets ยังไง", "least privilege คืออะไร", "test pyramid", "จะ deploy อย่างปลอดภัยยังไง".

มองแบบ SRE: security กับ testing คือ กลไกลด incident — least privilege จำกัด blast radius, และ test/canary จับปัญหาก่อนถึงผู้ใช้.

Core Concepts — Security

Least Privilege & IAM

ให้สิทธิ์ น้อยที่สุดที่จำเป็น ต่อ identity/service. จำกัด blast radius เมื่อ credential รั่วหรือ service ถูก compromise. หลีกเลี่ยง wildcard *:* ในนโยบาย.

Secrets Management

  • อย่า hardcode secret ในโค้ด, env var ที่ commit, หรือ image layer.
  • เก็บใน secret store (AWS Secrets Manager / SSM Parameter Store SecureString / Vault) แล้ว inject ตอน runtime.
  • Rotate เป็นระยะ และเมื่อสงสัยว่ารั่ว; ใช้ short-lived credential (เช่น IRSA/STS) แทน long-lived key.

กับดักที่ทำให้ตกสัมภาษณ์ (และทำระบบพังจริง)

Secret ที่เผลอ commit ลง git จะอยู่ใน history แม้จะลบไฟล์แล้ว — ต้องถือว่า "รั่วแล้ว" และ rotate ทันที ไม่ใช่แค่ลบ commit. เช่นเดียวกับ secret ที่ฝังใน image layer (ดูด้วย docker history ได้).

Network Security

  • Segmentation / least exposure: เปิด port/route เท่าที่จำเป็น. Security Group แบบ allow-list, ปิด public ที่ไม่ต้อง.
  • Kubernetes NetworkPolicy: default คือ pod คุยกันได้หมด — ตั้ง policy จำกัดว่า pod ไหน คุยกับใครได้ (เช่น เฉพาะ app → db).
  • Zero trust: อย่าเชื่อแค่เพราะอยู่ใน network เดียวกัน — ยัง authenticate/authorize ทุก call.

Encryption & Defense in Depth

  • In transit (TLS) และ at rest (disk/DB encryption, KMS).
  • Defense in depth = หลายชั้นป้องกัน (network + identity + app + data) เพื่อไม่ให้ชั้นเดียวพังแล้วจบ.

Supply chain

scan image หา CVE (เช่น Trivy, ECR scanning), pin base image, ใช้ image เล็ก (distroless) เพื่อลด attack surface.

Core Concepts — Testing

Test Pyramid

ยิ่งขึ้นสูง ยิ่งช้า/แพง/เปราะ — ฐานควรใหญ่ (unit เยอะ)
  • Unit — ทดสอบฟังก์ชันเดี่ยว เร็ว รันเยอะ ๆ ได้ (ฐานพีระมิด).
  • Integration — ทดสอบหลาย component ทำงานร่วมกัน (เช่น service + DB).
  • E2E — ทดสอบทั้ง flow เหมือนผู้ใช้จริง ช้าและเปราะ จึงมีน้อย.

Testing สำหรับ infra/reliability (มุม SRE)

  • Smoke test หลัง deploy — เช็กว่า critical path ยังทำงาน.
  • Canary / progressive rollout — ปล่อยเวอร์ชันใหม่ให้ traffic ส่วนน้อยก่อน สังเกต metric แล้วค่อยขยาย; พร้อม auto-rollback ถ้า error พุ่ง.
  • Load / stress test — หา breaking point และ capacity.
  • Chaos engineering — จงใจฉีด failure (kill pod, เพิ่ม latency) เพื่อพิสูจน์ว่าระบบทน ตามที่ออกแบบ (เช่น circuit breaker ทำงานจริง).
  • Health checks / probes — ให้ระบบ (LB/K8s) รู้ว่า instance พร้อมรับ traffic ไหม.

เชื่อม security + testing เข้ากับ reliability

เวลาโดนถาม ให้ผูกกลับไปที่ incident เสมอ: "least privilege ลด blast radius, canary + auto-rollback ลด MTTR และ blast radius ของ bad deploy, chaos engineering พิสูจน์ resilience pattern ก่อนของจริงจะพัง" — แสดงว่าคุณคิดเชิงระบบ ไม่ได้ท่องนิยาม.

เชื่อมกับ AWS / Cloud ที่คุณรู้อยู่แล้ว

แนวคิดบริการ AWS
Least privilege / identityIAM policy (จำกัด action/resource), IRSA บน EKS, STS short-lived creds
SecretsSecrets Manager, SSM Parameter Store (SecureString) — rotate ได้
EncryptionKMS (at rest), ACM/TLS (in transit)
NetworkSecurity Group/NACL, VPC, PrivateLink, K8s NetworkPolicy
Threat detection / scanningGuardDuty, Inspector, ECR image scanning
Safe deployCodeDeploy canary/linear + automatic rollback, CloudWatch alarms เป็น gate

คำถาม interview ที่เจอบ่อย + แนวคำตอบ

  1. "จัดการ secrets ใน production ยังไง?" → เก็บใน secret store (Secrets Manager/SSM/Vault), inject ตอน runtime, ไม่ commit/ฝัง image, rotate สม่ำเสมอ, ใช้ short-lived creds (IRSA/STS). ถ้าเผลอ leak → rotate ทันที.

  2. "least privilege คืออะไร ทำไมสำคัญ?" → ให้สิทธิ์เท่าที่จำเป็น เพื่อจำกัด blast radius เมื่อถูก compromise. ตัวอย่าง: service ที่ อ่าน S3 bucket เดียวควรมี policy เฉพาะ s3:GetObject บน bucket นั้น ไม่ใช่ s3:*.

  3. "อธิบาย test pyramid" → unit เยอะ (เร็ว/ถูก) เป็นฐาน, integration ปานกลาง, e2e น้อย (ช้า/เปราะ). โครงนี้ให้ feedback เร็วและ suite เสถียร.

  4. "จะ deploy เวอร์ชันใหม่อย่างปลอดภัยยังไง?" → canary/progressive rollout: ปล่อย traffic ส่วนน้อย, เฝ้า metric (error/latency), auto-rollback ถ้าเกินเกณฑ์; มี smoke test หลัง deploy. ลด blast radius และ MTTR.

  5. "chaos engineering คืออะไร มีประโยชน์ยังไง?" → จงใจฉีด failure ในสภาพควบคุมได้เพื่อพิสูจน์ว่าระบบทนจริง (timeout/circuit breaker/failover ทำงาน) — เจอจุดอ่อนก่อนที่ของจริงจะพังตอนตี 3.

Pitfalls — จุดที่ผู้สมัครมักพลาด

ระวังกับดักเหล่านี้

  • Hardcode secret / commit key ลง git — และคิดว่าลบ commit แล้วปลอดภัย (ต้อง rotate).
  • IAM policy กว้างเกิน (*:*) — blast radius มหาศาลเมื่อรั่ว.
  • ไม่ rotate secret / ใช้ long-lived key ทั้งที่ควรใช้ short-lived (STS/IRSA).
  • ทดสอบแค่ happy path — ไม่เทสต์ error/timeout/failure ที่เป็นตัวก่อ incident.
  • Deploy ตรงเข้า 100% โดยไม่มี canary/rollback — bad deploy กระทบทุกคนทันที.
  • มอง 'อยู่ใน VPC เดียวกัน = ปลอดภัย' — ละเลย zero trust / NetworkPolicy.

Quiz ท้ายบท

Quiz ท้ายบท

ตอบแล้ว 0/7
  1. 1.เผลอ commit AWS access key ลง git แล้วลบ commit นั้นทีหลัง ควรทำอย่างไร?

  2. 2.หลักการ least privilege หมายถึงอะไร?

  3. 3.ตาม test pyramid ชนิดของ test ที่ควรมี 'มากที่สุด' คือข้อใด?

  4. 4.วิธี deploy เวอร์ชันใหม่ที่ลด blast radius ของ bad deploy ได้ดีที่สุดคือข้อใด?

  5. 5.โดย default pod ใน Kubernetes สื่อสารกันได้อย่างไร และ NetworkPolicy ช่วยอะไร?

  6. 6.chaos engineering มีเป้าหมายหลักคืออะไร?

  7. 7.ทำไมการทดสอบแค่ 'happy path' จึงเป็นความเสี่ยงด้าน reliability?

Cheat Sheet — อ่านก่อนเข้าห้องสัมภาษณ์

สรุปเร็ว 30 วินาที

  • Least privilege = สิทธิ์น้อยสุดที่จำเป็น → จำกัด blast radius (เลี่ยง *:*)
  • Secrets: อย่า hardcode/commit/ฝัง image; ใช้ secret store + rotate + short-lived (IRSA/STS); leak แล้ว = rotate ทันที
  • Network: segmentation, K8s NetworkPolicy (default เปิดหมด), zero trust
  • Encryption: in transit (TLS) + at rest (KMS); defense in depth
  • Test pyramid: unit (เยอะ) → integration → e2e (น้อย)
  • Safe deploy: canary + เฝ้า metric + auto-rollback + smoke test; chaos พิสูจน์ resilience
  • ผูกทุกอย่างกลับไปที่ ลด blast radius / ลด MTTR
อ่านจบแล้ว? ทำเครื่องหมายไว้เพื่อติดตามความคืบหน้า