Practical Ways to Implement SRE
Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, and Stephen Thorne

#SRE
#Cloud
#monitor
#New_York_Times
📘 در سال ۲۰۱۶، کتاب مهندسی قابلیت اطمینان سایت (Site Reliability Engineering) گوگل یک بحث بزرگ در صنعت راه انداخت دربارهی اینکه اجرای سرویسهای پروداکشن در دنیای امروز واقعاً یعنی چی — و چرا موضوع پایداری، بخش بنیادی طراحی سرویسهاست. حالا مهندسهای گوگل که روی اون کتاب پرفروش کار کرده بودن، کتاب The Site Reliability Workbook رو معرفی میکنن؛ یک همراه عملی و Hands-on که با مثالهای واقعی نشون میده چطور اصول و Practiceهای SRE رو داخل محیط خودت پیادهسازی کنی.
🛠️ این ورکبوک جدید فقط مثالهای عملی از تجربههای گوگل رو جمع نکرده، بلکه Case Studyهایی از مشتریهای Google Cloud Platform هم ارائه میده که این مسیر رو طی کردن. شرکتهایی مثل:
و شرکتهای دیگه، تجربههای واقعی و سختیهایی که پشت سر گذاشتن رو توضیح میدن؛ اینکه چه چیزهایی براشون جواب داده و چه چیزهایی شکست خورده.
🚀 وارد این ورکبوک میشی و یاد میگیری چطور Practice مخصوص SRE خودت رو توسعه بدی؛ فرقی هم نمیکنه شرکتت کوچک باشه یا بزرگ.
🔥 چیزهایی که یاد میگیری:
🔹 چطور سرویسهای پایدار رو در محیطهایی اجرا کنی که کنترل کامل روی اونها نداری — مثل فضای ابری
🔹 کاربردهای عملی برای ساخت، مانیتورینگ و اجرای سرویسها با استفاده از اهداف سطح سرویس (Service Level Objectives یا SLO)
🔹 چطور تیمهای عملیاتی فعلی رو به تیم SRE تبدیل کنی — از جمله اینکه چطور از فشار و Overload عملیاتی خارج بشی
🔹 روشهای شروع SRE چه در پروژههای Greenfield و چه در سیستمهای Brownfield
📚 فهرست مطالب
فصل 1. ارتباط SRE با DevOps
بخش اول: پایهها
فصل 2. پیادهسازی SLOها
فصل 3. مطالعههای موردی مهندسی SLO
فصل 4. مانیتورینگ
فصل 5. هشداردهی بر اساس SLOها
فصل 6. حذف Toil
فصل 7. سادگی
بخش دوم: Practiceها
فصل 8. آنکال
فصل 9. پاسخ به Incident
فصل 10. فرهنگ پستمرتم: یادگیری از Failure
فصل 11. مدیریت بار
فصل 12. معرفی طراحی سیستمهای بزرگ بهصورت غیرانتزاعی
فصل 13. پایپلاینهای پردازش داده
فصل 14. طراحی پیکربندی و Best Practiceها
فصل 15. جزئیات پیکربندی
فصل 16. انتشار Canary
بخش سوم: فرایندها
فصل 17. شناسایی و بازیابی از Overload
فصل 18. مدل تعامل SRE
فصل 19. SRE: فراتر از مرزهای سازمان
فصل 20. چرخهی عمر تیمهای SRE
فصل 21. مدیریت تغییرات سازمانی در SRE
👤 درباره نویسندهها
✍️ بتسی بایر نویسندهی فنی تیم Google Site Reliability Engineering در نیویورک هست. او قبلاً مستندات تیمهای دیتاسنتر و عملیات سختافزار گوگل رو نوشته. قبل از مهاجرت به نیویورک هم مدرس نگارش فنی در دانشگاه استنفورد بوده.
🌍 نیال ریچارد مورفی در حال حاضر مدیر جهانی Azure SRE در مایکروسافت هست و در دفتر دوبلین ایرلند کار میکنه. او بیشتر از بیست سال در حوزهی زیرساخت اینترنت فعالیت داشته و مدرکهایی در رشتههای علوم کامپیوتر، ریاضیات و مطالعات شعر داره.
In 2016, Google’s Site Reliability Engineering book ignited an industry discussion on what it means to run production services today—and why reliability considerations are fundamental to service design. Now, Google engineers who worked on that bestseller introduce The Site Reliability Workbook, a hands-on companion that uses concrete examples to show you how to put SRE principles and practices to work in your environment.
This new workbook not only combines practical examples from Google’s experiences, but also provides case studies from Google’s Cloud Platform customers who underwent this journey. Evernote, The Home Depot, The New York Times, and other companies outline hard-won experiences of what worked for them and what didn’t.
Dive into this workbook and learn how to flesh out your own SRE practice, no matter what size your company is.
You’ll learn:
Table of Contents
Chapter 1. How SRE Relates to DevOps
Part I. Foundations
Chapter 2. Implementing SLOs
Chapter 3. SLO Engineering Case Studies
Chapter 4. Monitoring
Chapter 5. Alerting on SLOs
Chapter 6. Eliminating Toil
Chapter 7. Simplicity
Part II. Practices
Chapter 8. On-Call
Chapter 9. Incident Response
Chapter 10. Postmortem Culture: Learning from Failure
Chapter 11. Managing Load
Chapter 12. Introducing Non-Abstract Large System Design
Chapter 13. Data Processing Pipelines
Chapter 14. Configuration Design and Best Practices
Chapter 15. Configuration Specifics
Chapter 16. Canarying Releases
Part III. Processes
Chapter 17. Identifying and Recovering from Overload
Chapter 18. SRE Engagement Model
Chapter 19. SRE: Reaching Beyond Your Walls
Chapter 20. SRE Team Lifecycles
Chapter 21. Organizational Change Management in SRE
Conclusion
Appendix A. Example SLO Document
Appendix B. Example Error Budget Policy
Appendix C. Results of Postmortem Analysis
Betsy Beyer is a Technical Writer for Google Site Reliability Engineering in NYC. She has previously written documentation for Google Datacenters and Hardware Operations teams. Before moving to New York, Betsy was a lecturer on technical writing at Stanford University.
Niall Richard Murphy is currently the global head of Azure SRE at Microsoft, working in their Dublin, Ireland office. He has worked in Internet infrastructure for over twenty years, and holds degrees in Computer Science, Mathematics, and Poetry Studies.









