0
نام کتاب
Site Reliability Engineering

How Google Runs Production Systems

Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy

Print Length550 Pages
PublisherO'Reilly
Edition1
LanguageEnglish
Year2016
ISBN9781491929124
1K
A1260
انتخاب نوع چاپ:
جلد سخت
1,405,000ت
0
جلد نرم
1,505,000ت(2 جلدی)
0
طلق پاپکو و فنر
1,545,000ت(2 جلدی)
0
مجموع:
0تومان
کیفیت متن:اورجینال انتشارات
قطع:B5
رنگ صفحات:دارای متن و کادر رنگی
پشتیبانی در روزهای تعطیل!
ارسال به سراسر کشور

#SRE

#Google

#software_engineer

#Monitoring

#Data

توضیحات

📘 بخش عظیمی از عمر یک سیستم نرم‌افزاری، نه صرف طراحی و پیاده‌سازی، بلکه صرف استفاده از اون میشه. پس چرا نگاه سنتی هنوز اصرار داره که مهندس‌های نرم‌افزار باید بیشتر تمرکزشون روی طراحی و توسعه‌ی سیستم‌های محاسباتی بزرگ‌مقیاس باشه؟


🧠 در این مجموعه از مقاله‌ها و Essayها، اعضای کلیدی تیم Site Reliability Engineering گوگل توضیح میدن چطور و چرا تمرکز روی کل چرخه‌ی عمر سیستم باعث شده گوگل بتونه بعضی از بزرگ‌ترین سیستم‌های نرم‌افزاری دنیا رو با موفقیت بسازه، دیپلوی کنه، مانیتور کنه و نگهداری کنه. توی این کتاب، اصول و Practiceهایی رو یاد میگیری که به مهندس‌های گوگل کمک میکنن سیستم‌ها رو مقیاس‌پذیرتر، پایدارتر و بهینه‌تر کنن — درس‌هایی که مستقیماً توی سازمان خودت هم قابل‌استفاده‌ست.


📚 این کتاب به چهار بخش تقسیم شده:

🔹 مقدمه — یاد میگیری مهندسی قابلیت اطمینان سایت (Site Reliability Engineering یا SRE) چیه و چرا با Practiceهای سنتی صنعت IT فرق داره

🔹 اصول — الگوها، رفتارها و حوزه‌هایی که روی کار یک مهندس SRE تأثیر میذارن رو بررسی میکنی

🔹 Practiceها — تئوری و کار عملی روزمره‌ی یک SRE رو درک میکنی؛ از ساخت تا اجرای سیستم‌های محاسباتی توزیع‌شده‌ی بزرگ

🔹 مدیریت — Best Practiceهای گوگل برای آموزش، ارتباطات و جلسه‌ها رو بررسی میکنی که سازمانت هم میتونه ازشون استفاده کنه


📖 چطور این کتاب رو بخونی

این کتاب مجموعه‌ای از مقاله‌هاست که توسط اعضا و فارغ‌التحصیل‌های سازمان Google Site Reliability Engineering نوشته شده. ساختارش بیشتر شبیه مجموعه‌مقاله‌های کنفرانسیه تا یک کتاب معمولی که توسط یک یا چند نویسنده نوشته شده باشه.

هر فصل طوری نوشته شده که بخشی از یک مجموعه‌ی منسجم باشه، اما اگر فقط موضوع خاصی برات جذابه، میتونی همون بخش رو جداگانه بخونی و باز هم چیز زیادی یاد بگیری. اگر مقاله‌های دیگه‌ای وجود داشته باشن که متن رو کامل‌تر کنن، داخل کتاب بهشون ارجاع داده شده تا بتونی ادامه بدی.


🛠️ لازم نیست کتاب رو با ترتیب خاصی بخونی، ولی پیشنهاد میشه حداقل از فصل‌های ۲ و ۳ شروع کنی؛ فصل‌هایی که به‌ترتیب محیط پروداکشن گوگل و نحوه‌ی برخورد SRE با ریسک رو توضیح میدن. در خیلی از جهات، ریسک مهم‌ترین ویژگی شغل ماست.

البته خوندن کتاب از اول تا آخر هم کاملاً ممکن و مفیده. فصل‌ها به‌صورت موضوعی دسته‌بندی شدن:

  • اصول
  • Practiceها
  • مدیریت

هر بخش هم مقدمه‌ی کوتاهی داره که توضیح میده مقاله‌های اون قسمت درباره‌ی چی هستن و به مقاله‌های دیگه‌ای از مهندس‌های SRE گوگل ارجاع میده که موضوعات خاص رو عمیق‌تر پوشش میدن. علاوه بر این، داخل کتاب به یک وب‌سایت همراه هم اشاره شده که منابع مفیدی ارائه میکنه.


💬 امیدواریم این کتاب حداقل به‌اندازه‌ای که ساختنش برای ما مفید و جذاب بود، برای شما هم مفید و جذاب باشه.

— ویراستارها


📚 این کتاب به چهار بخش تقسیم شده:

🔹 مقدمه — یاد میگیری مهندسی قابلیت اطمینان سایت چیه و چرا با Practiceهای سنتی صنعت IT فرق داره

🔹 اصول — الگوها، رفتارها و دغدغه‌هایی که روی کار مهندس SRE تأثیر میذارن رو بررسی میکنی

🔹 Practiceها — تئوری و کار عملی روزمره‌ی SREها در ساخت و اجرای سیستم‌های محاسباتی توزیع‌شده‌ی بزرگ رو درک میکنی

🔹 مدیریت — Best Practiceهای گوگل برای آموزش، ارتباطات و جلسه‌ها رو بررسی میکنی که سازمانت میتونه ازشون استفاده کنه


📚 فهرست مطالب

بخش اول: مقدمه

1. مقدمه

2. محیط پروداکشن در گوگل از دیدگاه یک SRE


بخش دوم: اصول

3. پذیرفتن ریسک

4. اهداف سطح سرویس

5. حذف Toil

6. مانیتورینگ سیستم‌های توزیع‌شده

7. تکامل اتوماسیون در گوگل

8. مهندسی انتشار

9. سادگی


بخش سوم: Practiceها

10. هشداردهی عملی با استفاده از داده‌های سری زمانی

11. آن‌کال بودن

12. عیب‌یابی مؤثر

13. پاسخ اضطراری

14. مدیریت Incidentها

15. فرهنگ Postmortem: یادگیری از Failure

16. ردیابی قطعی‌ها

17. تست برای پایداری

18. مهندسی نرم‌افزار در SRE

19. توزیع بار در Frontend

20. توزیع بار در دیتاسنتر

21. مدیریت Overload

22. مقابله با Failureهای زنجیره‌ای

23. مدیریت وضعیت بحرانی: اجماع توزیع‌شده برای پایداری

24. زمان‌بندی دوره‌ای توزیع‌شده با Cron

25. پایپ‌لاین‌های پردازش داده

26. یکپارچگی داده: چیزی که میخوانی همان چیزی است که نوشته‌ای

27. لانچ پایدار محصول در مقیاس بزرگ


بخش چهارم: مدیریت

28. رساندن SREها به مرحله‌ی On-Call و فراتر از آن

29. مدیریت Interruptها

30. قرار دادن یک SRE برای بازیابی از Overload عملیاتی

31. ارتباطات و همکاری در SRE

32. مدل در حال تکامل تعامل SRE


بخش پنجم: نتیجه‌گیری

33. درس‌هایی از صنایع دیگر

34. نتیجه‌گیری


پیوست A. جدول Availability

پیوست B. مجموعه‌ای از Best Practiceها برای سرویس‌های پروداکشن

پیوست C. نمونه سند وضعیت Incident

پیوست D. نمونه Postmortem

پیوست E. چک‌لیست هماهنگی لانچ

پیوست F. نمونه صورت‌جلسه‌ی جلسات پروداکشن


👤 درباره نویسنده‌ها

🌍 نیال مورفی رهبری تیم Ads Site Reliability Engineering گوگل در ایرلند رو برعهده داره. حدود ۲۰ ساله که در صنعت اینترنت فعالیت میکنه و در حال حاضر رئیس INEX، هاب Peering ایرلند، هست. او نویسنده یا هم‌نویسنده‌ی چندین مقاله و کتاب فنی از جمله IPv6 Network Administration برای انتشارات O’Reilly و تعدادی RFC بوده. الان هم مشغول نوشتن تاریخچه‌ی اینترنت در ایرلنده. او مدرک‌هایی در علوم کامپیوتر، ریاضیات و مطالعات شعر داره؛ ترکیبی که به گفته‌ی خودش احتمالاً یک اشتباه عجیب بوده. او همراه همسر و دو پسرش در دوبلین زندگی میکنه.


✍️ بتسی بایر نویسنده‌ی فنی تیم Google Site Reliability Engineering در نیویورکه. او قبلاً مستندات تیم‌های دیتاسنتر و عملیات سخت‌افزار گوگل رو نوشته. قبل از رفتن به نیویورک هم مدرس نگارش فنی در دانشگاه استنفورد بوده.


☁️ کریس جونز مهندس SRE سرویس Google App Engine هست؛ پلتفرم ابری PaaS گوگل که روزانه بیشتر از ۲۸ میلیارد درخواست رو پردازش میکنه. او در سان‌فرانسیسکو مستقره و قبلاً مسئول نگهداری و مدیریت سیستم‌های آماری تبلیغات گوگل، انبار داده و سیستم‌های پشتیبانی مشتری بوده. در دوره‌های مختلف زندگیش، در IT دانشگاهی کار کرده، داده‌های کمپین‌های سیاسی رو تحلیل کرده و کمی هم روی کرنل BSD هک کرده. او در مسیرش مدرک‌هایی در مهندسی کامپیوتر، اقتصاد و سیاست‌گذاری فناوری گرفته و همچنین مهندس حرفه‌ای دارای مجوزه.


🧪 جنیفر پتف مدیر برنامه‌ی تیم Site Reliability Engineering گوگله و در دوبلین ایرلند مستقره. او پروژه‌های جهانی بزرگی رو در حوزه‌های مختلفی مثل تحقیقات علمی، مهندسی، منابع انسانی و عملیات تبلیغات مدیریت کرده. جنیفر بعد از هشت سال فعالیت در صنعت شیمی به گوگل پیوست. او دکترای شیمی از دانشگاه استنفورد و مدرک کارشناسی شیمی و روان‌شناسی از دانشگاه University of Rochester داره.


The overwhelming majority of a software system's lifespan is spent in use, not in design or implementation. So, why does conventional wisdom insist that software engineers focus primarily on the design and development of large-scale computing systems?


In this collection of essays and articles, key members of Google's Site Reliability Team explain how and why their commitment to the entire lifecycle has enabled the company to successfully build, deploy, monitor, and maintain some of the largest software systems in the world. You'll learn the principles and practices that enable Google engineers to make systems more scalable, reliable, and efficient―lessons directly applicable to your organization.


This book is divided into four sections:

  • Introduction―Learn what site reliability engineering is and why it differs from conventional IT industry practices
  • Principles―Examine the patterns, behaviors, and areas of concern that influence the work of a site reliability engineer (SRE)
  • Practices―Understand the theory and practice of an SRE's day-to-day work: building and operating large distributed computing systems
  • Management―Explore Google's best practices for training, communication, and meetings that your organization can use


How to Read This Book

This book is a series of essays written by members and alumni of Google’s Site Reliability Engineering organization. It’s much more like conference proceedings than it is like a standard book by an author or a small number of authors. Each chapter is intended to be read as a part of a coherent whole, but a good deal can be gained by reading on whatever subject particularly interests you. (If there are other articles that support or inform the text, we reference them so you can follow up accordingly.)


You don’t need to read in any particular order, though we’d suggest at least starting with Chapters 2 and 3, which describe Google’s production environment and outline how SRE approaches risk, respectively. (Risk is, in many ways, the key quality of our profession.) Reading cover-to-cover is, of course, also useful and possible; our chapters are grouped thematically, into Principles (Part II), Practices (Part III), and Management (Part IV). Each has a small introduction that highlights what the individual pieces are about, and references other articles published by Google SREs, covering specific topics in more detail. Additionally, there’s a companion website mentioned in the book that has a number of helpful resources.


We hope this will be at least as useful and interesting to you as putting it together was for us.

— The Editors.


This book is divided into four sections:
  • Introduction—Learn what site reliability engineering is and why it differs from conventional IT industry practices
  • Principles—Examine the patterns, behaviors, and areas of concern that influence the work of a site reliability engineer (SRE)
  • Practices—Understand the theory and practice of an SRE’s day-to-day work: building and operating large distributed computing systems
  • Management—Explore Google's best practices for training, communication, and meetings that your organization can use


Table of Contents

Part I. Introduction

  Chapter 1. Introduction

  Chapter 2. The Production Environment at Google, from the Viewpoint of an SRE

Part II. Principles

  Chapter 3. Embracing Risk

  Chapter 4. Service Level Objectives

  Chapter 5. Eliminating Toil

  Chapter 6. Monitoring Distributed Systems

  Chapter 7. The Evolution of Automation at Google

  Chapter 8. Release Engineering

  Chapter 9. Simplicity

Part III. Practices

  Chapter 10. Practical Alerting from Time-Series Data

  Chapter 11. Being On-Call

  Chapter 12. Effective Troubleshooting

  Chapter 13. Emergency Response

  Chapter 14. Managing Incidents

  Chapter 15. Postmortem Culture: Learning from Failure

  Chapter 16. Tracking Outages

  Chapter 17. Testing for Reliability

  Chapter 18. Software Engineering in SRE

  Chapter 19. Load Balancing at the Frontend

  Chapter 20. Load Balancing in the Datacenter

  Chapter 21. Handling Overload

  Chapter 22. Addressing Cascading Failures

  Chapter 23. Managing Critical State: Distributed Consensus for Reliability

  Chapter 24. Distributed Periodic Scheduling with Cron

  Chapter 25. Data Processing Pipelines

  Chapter 26. Data Integrity: What You Read Is What You Wrote

  Chapter 27. Reliable Product Launches at Scale

Part IV. Management

  Chapter 28. Accelerating SREs to On-Call and Beyond

  Chapter 29. Dealing with Interrupts

  Chapter 30. Embedding an SRE to Recover from Operational Overload

  Chapter 31. Communication and Collaboration in SRE

  Chapter 32. The Evolving SRE Engagement Model

Part V. Conclusions

  Chapter 33. Lessons Learned from Other Industries

  Chapter 34. Conclusion

Appendix A. Availability Table

Appendix B. A Collection of Best Practices for Production Services

Appendix C. Example Incident State Document

Appendix D. Example Postmortem

Appendix E. Launch Coordination Checklist

Appendix F. Example Production Meeting Minutes


About the Author

Niall Murphy leads the Ads Site Reliability Engineering team at Google Ireland. He has been involved in the Internet industry for about 20 years, and is currently chairperson of INEX, Ireland’s peering hub. He is the author or coauthor of a number of technical papers and/or books, including "IPv6 Network Administration" for O’Reilly, and a number of RFCs. He is currently cowriting a history of the Internet in Ireland, and is the holder of degrees in Computer Science, Mathematics, and Poetry Studies, which is surely some kind of mistake. He lives in Dublin with his wife and two sons.


Betsy Beyer is a Technical Writer for Google Site Reliability Engineering in NYC. She has previously written documentation for Google Datacenters and Hardware Operations teams. Before moving to New York, Betsy was a lecturer on technical writing at Stanford University.


Chris Jones is a Site Reliability Engineer for Google App Engine, a cloud platform-as-a-service product serving over 28 billion requests per day. Based in San Francisco, he has previously been responsible for the care and feeding of Google’s advertising statistics, data warehousing, and customer support systems. In other lives, Chris has worked in academic IT, analyzed data for political campaigns, and engaged in some light BSD kernel hacking, picking up degrees in Computer Engineering, Economics, and Technology Policy along the way. He’s also a licensed professional engineer.


Jennifer Petoff is a Program Manager for Google’s Site Reliability Engineering team and based in Dublin, Ireland. She has managed large global projects across wide-ranging domains including scientific research, engineering, human resources, and advertising operations. Jennifer joined Google after spending eight years in the chemical industry. She holds a PhD in Chemistry from Stanford University and a BS in Chemistry and a BA in Psychology from the University of Rochester.

دیدگاه خود را بنویسید
نظرات کاربران (0 دیدگاه)
نظری وجود ندارد.
کتاب های مشابه
Software Engineering
808
Reengineering Software
634,000 تومان
Software Engineering
1,100
Own Your Tech Career
624,000 تومان
Software Engineering
732
Latency
623,000 تومان
Software Engineering
1,085
Object-Oriented Software Engineering
1,298,000 تومان
Software Engineering
997
The Rational Software Engineer
572,000 تومان
Software Engineering
1,089
Team Geek
509,000 تومان
Software Engineering
1,195
Managing Humans
805,000 تومان
Software Engineering
779
The Effective Software Engineer
624,000 تومان
الگوریتم‌‌ها
1,205
Software Engineering and Algorithms
1,608,000 تومان
Java
646
Object-Oriented Software Engineering Using UML, Patterns, and Java
1,856,000 تومان
قیمت
منصفانه
ارسال به
سراسر کشور
تضمین
کیفیت
پشتیبانی در
روزهای تعطیل
خرید امن
و آسان
آرشیو بزرگ
کتاب‌های تخصصی
هـر روز با بهتــرین و جــدیــدتـرین
کتاب های روز دنیا با ما همراه باشید
آدرس
پشتیبانی
مدیریت
ساعات پاسخگویی
درباره اسکای بوک
دسترسی های سریع
  • راهنمای خرید
  • راهنمای ارسال
  • سوالات متداول
  • قوانین و مقررات
  • وبلاگ
  • درباره ما
چاپ دیجیتال اسکای بوک. 2024-2022 ©