Applying SRE Principles to ML in Production
Cathy Chen, Niall Richard Murphy, Kranti Parisa, D. Sculley, and Todd Underwood

#ML
#Machine_Learning
#SRE
🏢 چه در یک استارتاپ کوچک کار کنی و چه در یک شرکت چندملیتی، این کتاب عملی به دانشمندهای داده، مهندسهای نرمافزار و SRE، مدیرهای محصول و صاحبان کسبوکار نشون میده چطور یادگیری ماشین رو بهشکل قابلاعتماد، مؤثر و مسئولانه در سازمان خودشون راهاندازی و مدیریت کنن.
📊 از مانیتور کردن مدلها در پروداکشن گرفته تا ساخت و اداره یک تیم توسعه مدل هماهنگ و کارآمد در یک سازمان محصولمحور، این کتاب تقریباً همه جنبههای عملی اجرای ML رو پوشش میده.
🧠 نویسندهها و متخصصهای مهندسی یعنی کتی چن، کرانتی پاریسا، نایل ریچارد مورفی، دی. اسکالی، تاد آندروود و نویسندههای مهمان، با بهکارگیری ذهنیت SRE در یادگیری ماشین توضیح میدن چطور یک سیستم ML کارآمد و قابلاعتماد بسازی و اجرا کنی.
🎯 چه هدفت افزایش درآمد باشه، چه بهینه کردن تصمیمگیری، حل مسئله یا فهم و اثرگذاری روی رفتار مشتری، یاد میگیری کارهای روزمره ML رو انجام بدی و همزمان تصویر بزرگتر رو هم از دست ندی.
🔍 موضوعهایی که بررسی میکنی
🧩 یادگیری ماشین چیست، چطور کار میکنه و به چه چیزهایی وابسته است
🔄 فریمورکهای مفهومی برای درک نحوه کار چرخههای ML
🚀 اینکه عملیاتیسازی درست چطور سیستمهای ML رو بهسادگی قابلمانیتور، قابلدیپلوی و قابلاجرا میکنه
🛠️ اینکه چرا سیستمهای ML عیبیابی در پروداکشن رو سختتر میکنن و چطور باید این مشکل رو جبران کرد
🤝 اینکه تیمهای ML، محصول و پروداکشن چطور میتونن ارتباط مؤثرتری با هم داشته باشن
💬 نظرها
💭 «یک بررسی عمیق و مستقل از مدل درباره جنبههای محصولی و فنی سیستمهای ML. راهنمایی که هر تیمی برای شناسایی و مدیریت Incidentها در مسیر رسیدن به Reliability باید داشته باشه.»
—گوکو موهانداس، بنیانگذار Made With ML
💭 «تخصص خودت در یادگیری ماشین رو تقویت کردی و حالا آمادهای ایدههات رو وارد پروداکشن کنی. این گنجینه از توصیههای متخصصهای باتجربه کمک میکنه این مسیر روانتر طی بشه و همزمان، ملاحظات اخلاقی و سازمانی مهم رو هم بهت یادآوری میکنه.»
—دیوید جی. گروم
💭 «کتاب Reliable Machine Learning برای هر کسی که سیستمهای واقعی یادگیری ماشین میسازه، ضروریه. این کتاب یک Blueprint برای فکر کردن درباره مسئلههای پیچیده و ظریف توسعه محصولهای مجهز به ML ارائه میده.»
—برایان اسپیرینگ، مدرس Data Science
💭 «در دنیایی که ML به بخشی از رویکرد پیشفرض برای حل مسئلهها تبدیل شده، ساخت راهکاری قابلاعتماد و مقیاسپذیر دیگه یک انتخاب نیست، بلکه ضرورته. این کتاب پایههای ساخت یک سیستم ML قابلاتکا رو فراهم میکنه.»
—جیمز بلسینگ
💭 «مهم نیست قبلاً چقدر Data Science کار کردی یا چقدر روی مبانی آماری یادگیری ماشین مسلطی. مهم نیست تمام کد منبع TensorFlow رو خطبهخط خوندی یا آموزش توزیعشده ML رو از صفر پیادهسازی کردی.
قبل از اینکه هر سیستم واقعی مبتنی بر یادگیری ماشین رو دیپلوی کنی، خوندن این کتاب برات مفیده. این دقیقاً همون چیزیه که هزاران دیپلوی آینده ML بهش نیاز دارن؛ دیپلویهایی که کاربردشون مثل یک شمشیر دولبه است. هرچقدر سیستم مفیدتر باشه، ریسکهای مربوط به امنیت، ایمنی، مشتریهای پرداختکننده، عدالت و تصمیمهای سیاستی هم بیشتر میشن.
این کتاب تمام عملیاتهایی رو که در چنین سطحی از مسئولیت باید اجرا کنی، بهطور کامل بررسی میکنه و میتونی مطمئن باشی که این مطالب حاصل دههها تجربه سخت و واقعی هستن.»
—اندرو مور، معاون Google
📖 فهرست مطالب
فصل ۱. مقدمه
فصل ۲. اصول مدیریت داده
فصل ۳. مقدمهای پایه بر مدلها
فصل ۴. Featureها و دادههای آموزشی
فصل ۵. ارزیابی اعتبار و کیفیت مدل
فصل ۶. عدالت، حریم خصوصی و سیستمهای اخلاقی ML
فصل ۷. سیستمهای آموزش
فصل ۸. سروینگ
فصل ۹. مانیتورینگ و مشاهدهپذیری مدلها
فصل ۱۰. یادگیری ماشین پیوسته
فصل ۱۱. پاسخ به Incident
فصل ۱۲. تعامل محصول و ML
فصل ۱۳. یکپارچهسازی ML در سازمان
فصل ۱۴. مثالهای عملی پیادهسازی ساختار سازمانی ML
فصل ۱۵. مطالعههای موردی: MLOps در عمل
👤 درباره نویسندگان
👩💼 کتی چن، دارای گواهی CPCC و مدرک کارشناسی ارشد، در کوچینگ رهبرهای تکنولوژی و کمک به توسعه مهارتهای رهبری تیم تخصص داره. او سابقه فعالیت در نقشهای Technical Program Manager، Product Manager و Engineering Manager رو داره.
🏗️ کتی در شرکتهای بزرگ تکنولوژی و استارتاپها، تیمهایی رو برای عرضه قابلیتهای محصول، ابزارهای داخلی و مدیریت سیستمهای بزرگ رهبری کرده. او مدرک کارشناسی مهندسی برق از UC Berkeley و کارشناسی ارشد روانشناسی سازمانی از Teachers College دانشگاه Columbia داره.
🌍 کتی در حال حاضر در پیتسبورگ پنسیلوانیا زندگی میکنه و در تیم SRE شرکت Google فعالیت داره.
👨💻 نایل مورفی از اواسط دهه ۱۹۹۰ در حوزه زیرساخت اینترنت کار کرده و روی سرویسهای آنلاین بزرگ تخصص داره. او با تمام ارائهدهندههای اصلی Cloud از دفترهای دوبلین اونها همکاری داشته و آخرین نقش مهمش در Microsoft، مدیریت جهانی Azure Site Reliability Engineering بوده.
🤖 اولین تجربه جدی او با یادگیری ماشین زمانی شکل گرفت که تیمهای ML تبلیغات Google در دوبلین رو مدیریت میکرد و با تاد آندروود در پیتسبورگ همکاری داشت. از اون زمان، ML همچنان یکی از حوزههای موردعلاقه او باقی مونده.
📚 نایل آغازگر، همنویسنده و ویراستار دو کتاب SRE شرکت Google است و احتمالاً یکی از معدود افرادی در دنیاست که همزمان در علوم کامپیوتر، ریاضی و مطالعات شعر مدرک دانشگاهی داره. او در دوبلین با همسر و دو فرزندش زندگی میکنه و روی یک استارتاپ در زمینه استفاده از ML در فضای SRE کار میکنه.
👨💼 کرانتی کی. پاریسا در حال حاضر معاون و رئیس Product Engineering در Dialpad است. تیمهای او نرمافزارهای ارتباطی و همکاری بلادرنگ، Cloud-Native و بزرگمقیاس میسازن که از تکنولوژیهای داخلی پیشرفته AI/ML و تلفنی استفاده میکنن.
🔎 پیش از Dialpad، او تیمهای مسئول پلتفرمها، محصولها و سرویسهای Search و Personalization در Apple رو رهبری کرده. کرانتی همینطور همبنیانگذار، CTO و مشاور فنی چند استارتاپ در حوزه Cloud Computing، SaaS و Enterprise Search بوده.
🧩 او در کامیونیتی Apache Lucene/Solr مشارکت داشته و همنویسنده کتاب Apache Solr Enterprise Search Server است. دولت آمریکا هم بهدلیل مشارکتهای برجسته او در Search و Discovery، عنوان Person of Extraordinary Ability یا EB1A رو بهش اعطا کرده.
👨🔬 دی. اسکالی در حال حاضر مدیرعامل Kaggle و مدیر بخش اکوسیستمهای ML شخص ثالث در Google است. او قبلاً مدیر تیم Google Brain و مسئول بعضی از حیاتیترین پایپلاینهای یادگیری ماشین پروداکشن Google بوده.
🛠️ تمرکز اصلی او روی بدهی فنی در یادگیری ماشین، Robustness و Reliability مدلها و پایپلاینهاست. او تیمهایی رو رهبری کرده که ML رو روی مسئلههای بسیار متنوعی به کار گرفتهاند؛ از پیشبینی کلیک روی تبلیغات و جلوگیری از سوءاستفاده گرفته تا طراحی پروتئین و کشف علمی.
🎓 دی. اسکالی همینطور در ساخت دوره Google Machine Learning Crash Course نقش داشته؛ دورهای که میلیونها نفر در سراسر دنیا باهاش ML یاد گرفتهاند.
👨💼 تاد آندروود مدیر ارشد در Google و رهبر تیم Machine Learning SRE است. او همینطور مسئول دفتر Google در پیتسبورگه.
⚙️ تیمهای ML SRE سرویسهای داخلی و خارجی یادگیری ماشین رو میسازن و اسکیل میکنن و در تقریباً تمام محصولهای مهم Google نقش حیاتی دارن.
🌐 تاد قبل از Google در شرکت Renesys مسئول عملیات، امنیت و Peering سرویسهای اطلاعات اینترنتی بود؛ شرکتی که حالا بخشی از Oracle Cloud است. قبل از اون هم CTO شرکت Oso Grande، یک ارائهدهنده مستقل اینترنت در نیومکزیکو، بود.
Whether you're part of a small startup or a multinational corporation, this practical book shows data scientists, software and site reliability engineers, product managers, and business owners how to run and establish ML reliably, effectively, and accountably within your organization. You'll gain insight into everything from how to do model monitoring in production to how to run a well-tuned model development team in a product organization.
By applying an SRE mindset to machine learning, authors and engineering professionals Cathy Chen, Kranti Parisa, Niall Richard Murphy, D. Sculley, Todd Underwood, and featured guest authors show you how to run an efficient and reliable ML system. Whether you want to increase revenue, optimize decision making, solve problems, or understand and influence customer behavior, you'll learn how to perform day-to-day ML tasks while keeping the bigger picture in mind.
You'll examine:
"A great model-agnostic deep dive into the product and technical aspects of ML systems. A guide every team should have for identifying and managing incidents when striving for reliability." - Goku Mohandas, Founder of Made With ML
"You've honed your machine learning expertise and are ready for your ideas to enter production — this treasure trove of tips from experienced practitioners will help ensure that journey is a smooth one, while also highlighting important ethical and organisational considerations." - David J. Groom
"Reliable Machine Learning is a must-read for people building real-world machine learning systems. It provides a blueprint for thinking about the complex and nuanced issues of developing machine learning enabled products." - Brian Spiering Data Science Instructor
"In a world where ML is becoming part of the default approach to problems, building a reliable and scalable solution is becoming a necessity. This book provides the groundwork for building an ML system that you can rely on." - James Blessing
"I don't care how much data science work you've done in the past, or how expert you are on the statistical foundations of Machine Learning. I don't care if you have read every line of the Tensorflow Source Code, or implemented your own distributed ML training from scratch. Before you ever put a real system based on Machine Learning into deployment you will benefit from reading this book. This is what is needed for the thousands of upcoming ML deployments where their usefulness is a double-edged sword. The more useful, the higher the stakes around safety, security, paying customers who are counting on you, fairness, or policy decisions that will be made on the basis of your system. This book thoroughly surveys the operations you need to be running if you have this level of responsibility, and you can rest assured that it comes from combined decades of hard won experience." - Andrew Moore, VP Google
Table of Contents
Chapter 1. Introduction
Chapter 2. Data Management Principles
Chapter 3. Basic Introduction to Models
Chapter 4. Feature and Training Data
Chapter 5. Evaluating Model Validity and Quality
Chapter 6. Fairness, Privacy, and Ethical ML Systems
Chapter 7. Training Systems
Chapter 8. Serving
Chapter 9. Monitoring and Observability for Models
Chapter 10. Continuous ML
Chapter 11. Incident Response
Chapter 12. How Product and ML Interact
Chapter 13. Integrating ML into Your Organization
Chapter 14. Practical ML Org Implementation Examples
Chapter 15. Case Studies: MLOps in Practice
Cathy Chen, CPCC, MA specializes in coaching tech leaders enabling development of their own skills in leading teams. She has held the role of technical program manager, product manager, and engineering manager. She has led teams in large tech companies and startups launching product features, internal tools, and operating large systems. Cathy has a BS in Electrical Engineering from UC Berkeley & MA in Organizational Psychology from Teachers College at Columbia University. Currently, Cathy lives with her partner in Pittsburgh, PA and works at Google in SRE.
Niall Murphy has worked in Internet infrastructure since the mid-1990s, specializing in large online services. He has worked with all of the major cloud providers from their Dublin, Ireland offices, and most recently at Microsoft, where he was global head of Azure Site Reliability Engineering (SRE). His first exposure to machine learning came with managing the Ads ML teams in Google’s Dublin office and working with Todd Underwood in Pittsburgh, though it has continued to fascinate him since. He is the instigator, co-author, and editor of the two Google SRE books, and he is probably one of the few people in the world to hold degrees in Computer Science, Mathematics, and Poetry Studies. He lives in Dublin with his wife and two children, and works on a startup involving ML in the SRE space
Kranti K. Parisa is currently the Vice President & Head of Product Engineering at Dialpad. His teams build large scale, cloud native real-time business communications & collaboration software with industry leading in-house AI/ML & Telephony technology. Before Dialpad, he has led teams that are responsible for search and personalization platforms, products and services at Apple. Kranti was a cofounder, CTO and technical advisor of multiple start-ups focusing on cloud computing, SaaS, and enterprise search. He has contributed to the Apache Lucene/Solr community and co-authored the book Apache Solr Enterprise Search Server. For his outstanding contributions to Search & Discovery, U.S. Government has recognized him as a Person of Extraordinary Ability (EB1A).
D. Sculley is currently the CEO of Kaggle and GM of Third Party ML Ecosystems at Google, and previously has been a Director in the Google Brain Team and the lead of some of Google's most critical production machine learning pipelines. He has focused on issues of technical debt in machine learning, along with robustness and reliability of models and pipelines, and has led teams applying machine learning to problems as diverse as ad click through prediction and abuse prevention to protein design and scientific discovery. Additionally, he has helped to create Google's Machine Learning Crash Course, which has taught ML to millions of people worldwide.
Todd Underwood is a Senior Director at Google and leads Machine Learning SRE. He is also Site Lead for Google’s Pittsburgh office. ML SRE teams build and scale internal and external ML services, and are critical to almost every significant product at Google. Before working at Google, Todd held a variety of roles at Renesys (in charge of operations, security, and peering for Internet intelligence services) now part of Oracle's Cloud, and before that he was Chief Technology Officer of Oso Grande, an independent Internet service provider in New Mexico.









