How Google Runs Production Systems
Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy

#SRE
#software_engineer
#Monitoring
#Data
📘 بخش عظیمی از عمر یک سیستم نرمافزاری، نه صرف طراحی و پیادهسازی، بلکه صرف استفاده از اون میشه. پس چرا نگاه سنتی هنوز اصرار داره که مهندسهای نرمافزار باید بیشتر تمرکزشون روی طراحی و توسعهی سیستمهای محاسباتی بزرگمقیاس باشه؟
🧠 در این مجموعه از مقالهها و Essayها، اعضای کلیدی تیم Site Reliability Engineering گوگل توضیح میدن چطور و چرا تمرکز روی کل چرخهی عمر سیستم باعث شده گوگل بتونه بعضی از بزرگترین سیستمهای نرمافزاری دنیا رو با موفقیت بسازه، دیپلوی کنه، مانیتور کنه و نگهداری کنه. توی این کتاب، اصول و Practiceهایی رو یاد میگیری که به مهندسهای گوگل کمک میکنن سیستمها رو مقیاسپذیرتر، پایدارتر و بهینهتر کنن — درسهایی که مستقیماً توی سازمان خودت هم قابلاستفادهست.
📚 این کتاب به چهار بخش تقسیم شده:
🔹 مقدمه — یاد میگیری مهندسی قابلیت اطمینان سایت (Site Reliability Engineering یا SRE) چیه و چرا با Practiceهای سنتی صنعت IT فرق داره
🔹 اصول — الگوها، رفتارها و حوزههایی که روی کار یک مهندس SRE تأثیر میذارن رو بررسی میکنی
🔹 Practiceها — تئوری و کار عملی روزمرهی یک SRE رو درک میکنی؛ از ساخت تا اجرای سیستمهای محاسباتی توزیعشدهی بزرگ
🔹 مدیریت — Best Practiceهای گوگل برای آموزش، ارتباطات و جلسهها رو بررسی میکنی که سازمانت هم میتونه ازشون استفاده کنه
📖 چطور این کتاب رو بخونی
این کتاب مجموعهای از مقالههاست که توسط اعضا و فارغالتحصیلهای سازمان Google Site Reliability Engineering نوشته شده. ساختارش بیشتر شبیه مجموعهمقالههای کنفرانسیه تا یک کتاب معمولی که توسط یک یا چند نویسنده نوشته شده باشه.
هر فصل طوری نوشته شده که بخشی از یک مجموعهی منسجم باشه، اما اگر فقط موضوع خاصی برات جذابه، میتونی همون بخش رو جداگانه بخونی و باز هم چیز زیادی یاد بگیری. اگر مقالههای دیگهای وجود داشته باشن که متن رو کاملتر کنن، داخل کتاب بهشون ارجاع داده شده تا بتونی ادامه بدی.
🛠️ لازم نیست کتاب رو با ترتیب خاصی بخونی، ولی پیشنهاد میشه حداقل از فصلهای ۲ و ۳ شروع کنی؛ فصلهایی که بهترتیب محیط پروداکشن گوگل و نحوهی برخورد SRE با ریسک رو توضیح میدن. در خیلی از جهات، ریسک مهمترین ویژگی شغل ماست.
البته خوندن کتاب از اول تا آخر هم کاملاً ممکن و مفیده. فصلها بهصورت موضوعی دستهبندی شدن:
هر بخش هم مقدمهی کوتاهی داره که توضیح میده مقالههای اون قسمت دربارهی چی هستن و به مقالههای دیگهای از مهندسهای SRE گوگل ارجاع میده که موضوعات خاص رو عمیقتر پوشش میدن. علاوه بر این، داخل کتاب به یک وبسایت همراه هم اشاره شده که منابع مفیدی ارائه میکنه.
💬 امیدواریم این کتاب حداقل بهاندازهای که ساختنش برای ما مفید و جذاب بود، برای شما هم مفید و جذاب باشه.
— ویراستارها
📚 این کتاب به چهار بخش تقسیم شده:
🔹 مقدمه — یاد میگیری مهندسی قابلیت اطمینان سایت چیه و چرا با Practiceهای سنتی صنعت IT فرق داره
🔹 اصول — الگوها، رفتارها و دغدغههایی که روی کار مهندس SRE تأثیر میذارن رو بررسی میکنی
🔹 Practiceها — تئوری و کار عملی روزمرهی SREها در ساخت و اجرای سیستمهای محاسباتی توزیعشدهی بزرگ رو درک میکنی
🔹 مدیریت — Best Practiceهای گوگل برای آموزش، ارتباطات و جلسهها رو بررسی میکنی که سازمانت میتونه ازشون استفاده کنه
📚 فهرست مطالب
بخش اول: مقدمه
1. مقدمه
2. محیط پروداکشن در گوگل از دیدگاه یک SRE
بخش دوم: اصول
3. پذیرفتن ریسک
4. اهداف سطح سرویس
5. حذف Toil
6. مانیتورینگ سیستمهای توزیعشده
7. تکامل اتوماسیون در گوگل
8. مهندسی انتشار
9. سادگی
بخش سوم: Practiceها
10. هشداردهی عملی با استفاده از دادههای سری زمانی
11. آنکال بودن
12. عیبیابی مؤثر
13. پاسخ اضطراری
14. مدیریت Incidentها
15. فرهنگ Postmortem: یادگیری از Failure
16. ردیابی قطعیها
17. تست برای پایداری
18. مهندسی نرمافزار در SRE
19. توزیع بار در Frontend
20. توزیع بار در دیتاسنتر
21. مدیریت Overload
22. مقابله با Failureهای زنجیرهای
23. مدیریت وضعیت بحرانی: اجماع توزیعشده برای پایداری
24. زمانبندی دورهای توزیعشده با Cron
25. پایپلاینهای پردازش داده
26. یکپارچگی داده: چیزی که میخوانی همان چیزی است که نوشتهای
27. لانچ پایدار محصول در مقیاس بزرگ
بخش چهارم: مدیریت
28. رساندن SREها به مرحلهی On-Call و فراتر از آن
29. مدیریت Interruptها
30. قرار دادن یک SRE برای بازیابی از Overload عملیاتی
31. ارتباطات و همکاری در SRE
32. مدل در حال تکامل تعامل SRE
بخش پنجم: نتیجهگیری
33. درسهایی از صنایع دیگر
34. نتیجهگیری
پیوست A. جدول Availability
پیوست B. مجموعهای از Best Practiceها برای سرویسهای پروداکشن
پیوست C. نمونه سند وضعیت Incident
پیوست D. نمونه Postmortem
پیوست E. چکلیست هماهنگی لانچ
پیوست F. نمونه صورتجلسهی جلسات پروداکشن
👤 درباره نویسندهها
🌍 نیال مورفی رهبری تیم Ads Site Reliability Engineering گوگل در ایرلند رو برعهده داره. حدود ۲۰ ساله که در صنعت اینترنت فعالیت میکنه و در حال حاضر رئیس INEX، هاب Peering ایرلند، هست. او نویسنده یا همنویسندهی چندین مقاله و کتاب فنی از جمله IPv6 Network Administration برای انتشارات O’Reilly و تعدادی RFC بوده. الان هم مشغول نوشتن تاریخچهی اینترنت در ایرلنده. او مدرکهایی در علوم کامپیوتر، ریاضیات و مطالعات شعر داره؛ ترکیبی که به گفتهی خودش احتمالاً یک اشتباه عجیب بوده. او همراه همسر و دو پسرش در دوبلین زندگی میکنه.
✍️ بتسی بایر نویسندهی فنی تیم Google Site Reliability Engineering در نیویورکه. او قبلاً مستندات تیمهای دیتاسنتر و عملیات سختافزار گوگل رو نوشته. قبل از رفتن به نیویورک هم مدرس نگارش فنی در دانشگاه استنفورد بوده.
☁️ کریس جونز مهندس SRE سرویس Google App Engine هست؛ پلتفرم ابری PaaS گوگل که روزانه بیشتر از ۲۸ میلیارد درخواست رو پردازش میکنه. او در سانفرانسیسکو مستقره و قبلاً مسئول نگهداری و مدیریت سیستمهای آماری تبلیغات گوگل، انبار داده و سیستمهای پشتیبانی مشتری بوده. در دورههای مختلف زندگیش، در IT دانشگاهی کار کرده، دادههای کمپینهای سیاسی رو تحلیل کرده و کمی هم روی کرنل BSD هک کرده. او در مسیرش مدرکهایی در مهندسی کامپیوتر، اقتصاد و سیاستگذاری فناوری گرفته و همچنین مهندس حرفهای دارای مجوزه.
🧪 جنیفر پتف مدیر برنامهی تیم Site Reliability Engineering گوگله و در دوبلین ایرلند مستقره. او پروژههای جهانی بزرگی رو در حوزههای مختلفی مثل تحقیقات علمی، مهندسی، منابع انسانی و عملیات تبلیغات مدیریت کرده. جنیفر بعد از هشت سال فعالیت در صنعت شیمی به گوگل پیوست. او دکترای شیمی از دانشگاه استنفورد و مدرک کارشناسی شیمی و روانشناسی از دانشگاه University of Rochester داره.
The overwhelming majority of a software system's lifespan is spent in use, not in design or implementation. So, why does conventional wisdom insist that software engineers focus primarily on the design and development of large-scale computing systems?
In this collection of essays and articles, key members of Google's Site Reliability Team explain how and why their commitment to the entire lifecycle has enabled the company to successfully build, deploy, monitor, and maintain some of the largest software systems in the world. You'll learn the principles and practices that enable Google engineers to make systems more scalable, reliable, and efficient―lessons directly applicable to your organization.
This book is divided into four sections:
This book is a series of essays written by members and alumni of Google’s Site Reliability Engineering organization. It’s much more like conference proceedings than it is like a standard book by an author or a small number of authors. Each chapter is intended to be read as a part of a coherent whole, but a good deal can be gained by reading on whatever subject particularly interests you. (If there are other articles that support or inform the text, we reference them so you can follow up accordingly.)
You don’t need to read in any particular order, though we’d suggest at least starting with Chapters 2 and 3, which describe Google’s production environment and outline how SRE approaches risk, respectively. (Risk is, in many ways, the key quality of our profession.) Reading cover-to-cover is, of course, also useful and possible; our chapters are grouped thematically, into Principles (Part II), Practices (Part III), and Management (Part IV). Each has a small introduction that highlights what the individual pieces are about, and references other articles published by Google SREs, covering specific topics in more detail. Additionally, there’s a companion website mentioned in the book that has a number of helpful resources.
We hope this will be at least as useful and interesting to you as putting it together was for us.
— The Editors.
Table of Contents
Part I. Introduction
Chapter 1. Introduction
Chapter 2. The Production Environment at Google, from the Viewpoint of an SRE
Part II. Principles
Chapter 3. Embracing Risk
Chapter 4. Service Level Objectives
Chapter 5. Eliminating Toil
Chapter 6. Monitoring Distributed Systems
Chapter 7. The Evolution of Automation at Google
Chapter 8. Release Engineering
Chapter 9. Simplicity
Part III. Practices
Chapter 10. Practical Alerting from Time-Series Data
Chapter 11. Being On-Call
Chapter 12. Effective Troubleshooting
Chapter 13. Emergency Response
Chapter 14. Managing Incidents
Chapter 15. Postmortem Culture: Learning from Failure
Chapter 16. Tracking Outages
Chapter 17. Testing for Reliability
Chapter 18. Software Engineering in SRE
Chapter 19. Load Balancing at the Frontend
Chapter 20. Load Balancing in the Datacenter
Chapter 21. Handling Overload
Chapter 22. Addressing Cascading Failures
Chapter 23. Managing Critical State: Distributed Consensus for Reliability
Chapter 24. Distributed Periodic Scheduling with Cron
Chapter 25. Data Processing Pipelines
Chapter 26. Data Integrity: What You Read Is What You Wrote
Chapter 27. Reliable Product Launches at Scale
Part IV. Management
Chapter 28. Accelerating SREs to On-Call and Beyond
Chapter 29. Dealing with Interrupts
Chapter 30. Embedding an SRE to Recover from Operational Overload
Chapter 31. Communication and Collaboration in SRE
Chapter 32. The Evolving SRE Engagement Model
Part V. Conclusions
Chapter 33. Lessons Learned from Other Industries
Chapter 34. Conclusion
Appendix A. Availability Table
Appendix B. A Collection of Best Practices for Production Services
Appendix C. Example Incident State Document
Appendix D. Example Postmortem
Appendix E. Launch Coordination Checklist
Appendix F. Example Production Meeting Minutes
Niall Murphy leads the Ads Site Reliability Engineering team at Google Ireland. He has been involved in the Internet industry for about 20 years, and is currently chairperson of INEX, Ireland’s peering hub. He is the author or coauthor of a number of technical papers and/or books, including "IPv6 Network Administration" for O’Reilly, and a number of RFCs. He is currently cowriting a history of the Internet in Ireland, and is the holder of degrees in Computer Science, Mathematics, and Poetry Studies, which is surely some kind of mistake. He lives in Dublin with his wife and two sons.
Betsy Beyer is a Technical Writer for Google Site Reliability Engineering in NYC. She has previously written documentation for Google Datacenters and Hardware Operations teams. Before moving to New York, Betsy was a lecturer on technical writing at Stanford University.
Chris Jones is a Site Reliability Engineer for Google App Engine, a cloud platform-as-a-service product serving over 28 billion requests per day. Based in San Francisco, he has previously been responsible for the care and feeding of Google’s advertising statistics, data warehousing, and customer support systems. In other lives, Chris has worked in academic IT, analyzed data for political campaigns, and engaged in some light BSD kernel hacking, picking up degrees in Computer Engineering, Economics, and Technology Policy along the way. He’s also a licensed professional engineer.
Jennifer Petoff is a Program Manager for Google’s Site Reliability Engineering team and based in Dublin, Ireland. She has managed large global projects across wide-ranging domains including scientific research, engineering, human resources, and advertising operations. Jennifer joined Google after spending eight years in the chemical industry. She holds a PhD in Chemistry from Stanford University and a BS in Chemistry and a BA in Psychology from the University of Rochester.









