4. Microservices — Design, Analysis & Troubleshooting
ทำไมหัวข้อนี้สำคัญกับ interview
microservices ทำให้ระบบ "พังเป็นบางส่วน" ได้ตลอดเวลา. ผู้สัมภาษณ์อยากรู้ว่าคุณเข้าใจ distributed-system failure modes และรู้จัก resilience pattern (timeout, retry, circuit breaker) ว่าใช้เมื่อไร/พังยังไง. คำถามคลาสสิก: "อธิบาย circuit breaker" และ "retry แบบไหนทำให้ระบบล่มหนักขึ้น".
กฎเหล็ก: ทุก network call ล้มเหลวได้ และ "partial failure" (บาง service ตอบ บางอันช้า) คือสภาพปกติ ไม่ใช่ข้อยกเว้น.
Core Concepts
Fallacies of Distributed Computing (ที่กระทบ SRE)
network ไม่ reliable, latency ไม่ เป็นศูนย์, bandwidth ไม่ ไม่จำกัด. ดังนั้น ต้องออกแบบเผื่อ timeout, retry, และ degradation เสมอ.
Timeout
ทุก call ต้องมี timeout ที่ชัดเจน. ถ้าไม่มี → thread/connection ค้างรอไม่จบ → resource หมด → ล่มลามทั้งระบบ. timeout ควรอิงจาก latency budget ไม่ใช่ตั้งมั่ว.
Retry (ต้องทำให้ถูก)
Retry ช่วยกับ error ชั่วคราว (transient) แต่ถ้าทำผิดจะกลายเป็น retry storm ที่ถล่ม service ที่กำลังป่วยให้ตายสนิท. กฎ:
- ใช้ exponential backoff + jitter (สุ่มหน่วง) เพื่อไม่ให้ทุก client retry พร้อมกัน.
- retry เฉพาะ operation ที่ idempotent (ทำซ้ำแล้วผลไม่เพี้ยน) หรือมี idempotency key.
- จำกัดจำนวนครั้ง + มี overall deadline.
retry delay = min(cap, base * 2^attempt) * random(0.5, 1.0) # backoff + jitter
เช่น base=100ms, cap=2s → ~100ms, 200ms, 400ms, ... (บวกสุ่ม)กับดักที่ถามบ่อยที่สุด
retry ไม่มี backoff/jitter + retry ทุกชั้น (client → gateway → service) = จำนวน request ทวีคูณตอนระบบกำลังป่วย → cascading failure. คำตอบที่ดีต้องพูดถึง jitter และ "อย่า retry ซ้อนหลายเลเยอร์".
Circuit Breaker
ป้องกันไม่ให้เรายิง request ไป dependency ที่กำลังพัง (fail fast แทนรอ timeout ทุกครั้ง). มี 3 สถานะ:
- Closed — ปกติ, ปล่อยผ่าน, นับ failure.
- Open — failure เกินเกณฑ์ → ตัดวงจร, request ล้มทันที (fail fast) ชั่วคราว.
- Half-open — หลังพักครบ ปล่อย request ทดลองจำนวนน้อย; ถ้าผ่านกลับ Closed, ถ้ายังพัง กลับ Open.
Cascading Failure & การป้องกัน
service A ช้า → caller ค้างรอ → thread เต็ม → caller ก็ช้า → ลามขึ้นไปเรื่อย ๆ. เครื่องมือกัน:
- Timeout สั้นและชัด (ตัดการรอที่ไม่จบ).
- Circuit breaker (fail fast จาก dependency ที่ป่วย).
- Bulkhead — แยก resource pool ต่อ dependency เพื่อไม่ให้ตัวที่ป่วยกินทรัพยากรหมด.
- Load shedding / rate limiting — ทิ้งงานเกินกำลังเพื่อรักษาส่วนที่เหลือ.
- Backpressure — ให้ผู้เรียกช้าลงแทนที่จะรับงานจนล้ม (เช่น queue มีขอบเขต).
- Graceful degradation — ตอบ partial/แคชแทนล่มทั้งหมด (เช่น ซ่อน recommendation แต่ยังซื้อได้).
Latency Budget
ถ้า API ต้องตอบใน 200ms และต้องเรียก 3 service ต่อเนื่อง งบเวลาต้องถูกแบ่ง (เช่น
80ms + 80ms + 40ms) รวม overhead. ช่วยตั้ง timeout ของแต่ละ hop อย่างมีเหตุผล และเห็นว่า
fan-out/serial call มากไปจะกินงบจนเกิน.
เชื่อมกับ AWS / Cloud ที่คุณรู้อยู่แล้ว
| Pattern | บริการ/กลไกบน AWS |
|---|---|
| Timeout/retry ที่ถูกต้อง | AWS SDK ทำ exponential backoff + jitter ให้อยู่แล้ว (ปรับ max attempts ได้) |
| Circuit breaker / mesh | AWS App Mesh (Envoy), API Gateway timeouts |
| Decoupling / backpressure | SQS (buffer งาน), Kinesis, EventBridge |
| Load shedding / throttling | API Gateway throttling, ALB, WAF rate rules |
| Retry orchestration | Step Functions (retry/catch ระดับ workflow) |
| Health-based routing | ALB/NLB health checks, Route 53 failover |
ผูกกับประสบการณ์ที่มี
ถ้าถูกถาม ให้ยกว่า "ใช้ SQS คั่นระหว่าง producer กับ consumer เพื่อดูดซับ spike และทำ backpressure ผ่าน queue depth" หรือ "ตั้ง API Gateway throttling กัน downstream ล้ม" — โยง pattern เข้ากับ service จริงที่คุณเคยใช้.
คำถาม interview ที่เจอบ่อย + แนวคำตอบ
-
"อธิบาย circuit breaker และ 3 สถานะ" → Closed (ปกติ), Open (fail fast เมื่อ failure เกินเกณฑ์), Half-open (ทดลองปล่อยน้อย ๆ เพื่อเช็กว่าหายหรือยัง). จุดประสงค์: ไม่ถล่ม dependency ที่ป่วยและไม่ค้างรอ timeout.
-
"retry แบบไหนทำให้ระบบล่มหนักขึ้น?" → retry ไม่มี backoff/jitter, retry ซ้อนหลายเลเยอร์, retry operation ที่ไม่ idempotent. แก้ด้วย exponential backoff + jitter, retry ชั้นเดียว, idempotency key, มี deadline.
-
"cascading failure เกิดยังไง ป้องกันยังไง?" → dependency ช้า → caller ค้างรอ → resource เต็ม → ลามต่อ. ป้องกันด้วย timeout, circuit breaker, bulkhead, load shedding, backpressure, graceful degradation.
-
"ทำไม timeout สำคัญมากใน microservices?" → ไม่มี timeout = รอไม่จบ = thread/connection หมด = ล่มลาม. timeout ควรอิง latency budget.
-
"idempotency คืออะไร เกี่ยวอะไรกับ retry?" → operation ที่ทำซ้ำแล้วผลเหมือนเดิม (เช่น PUT, หรือ POST ที่มี idempotency key). จำเป็นเพราะ retry อาจทำงานซ้ำ — ถ้าไม่ idempotent อาจตัดเงินซ้ำ.
Pitfalls — จุดที่ผู้สมัครมักพลาด
ระวังกับดักเหล่านี้
- Retry ไม่มี backoff/jitter → retry storm ถล่ม service ที่กำลังป่วย.
- ไม่มี timeout / timeout ยาวเกิน → รอค้างจน resource หมด.
- Retry operation ที่ไม่ idempotent → ทำงานซ้ำ (ตัดเงินซ้ำ, ส่งซ้ำ).
- ไม่มี circuit breaker / bulkhead → dependency เดียวป่วยลามทั้งระบบ.
- Fan-out / serial call มากเกินงบ latency → p99 พุ่งแม้แต่ละ hop ยังดู "ปกติ".
- มองว่า network เชื่อถือได้ → ไม่เผื่อ partial failure.
Quiz ท้ายบท
Quiz ท้ายบท
ตอบแล้ว 0/71.circuit breaker ในสถานะ 'Open' ทำอะไร?
2.อะไรคือสาเหตุหลักของ 'retry storm'?
3.ทำไม retry จึงควรทำเฉพาะกับ operation ที่ idempotent?
4.cascading failure ในระบบ microservices เกิดขึ้นได้อย่างไร?
5.pattern ใดช่วย 'แยก resource pool' เพื่อไม่ให้ dependency ที่ป่วยกินทรัพยากรของส่วนอื่นจนหมด?
6.API ต้องตอบใน 200ms แต่เรียก 3 downstream แบบต่อเนื่อง (serial) แนวคิด latency budget บอกอะไร?
7.ข้อใดคือตัวอย่างของ 'graceful degradation'?
Cheat Sheet — อ่านก่อนเข้าห้องสัมภาษณ์
สรุปเร็ว 30 วินาที
- ทุก network call ล้มเหลวได้ → ออกแบบเผื่อ partial failure เสมอ
- Timeout ทุก call (อิง latency budget) — ไม่มี timeout = resource หมด = ล่มลาม
- Retry = exponential backoff + jitter, ชั้นเดียว, เฉพาะ idempotent, มี deadline
- Circuit breaker: Closed → Open (fail fast) → Half-open (ทดลอง)
- กัน cascading: timeout · circuit breaker · bulkhead · load shedding · backpressure · graceful degradation
- Latency budget = แบ่งงบเวลาให้แต่ละ hop
- AWS: SDK backoff+jitter · App Mesh · SQS (backpressure) · API Gateway throttling · Step Functions retry