What every developer should know about large distributed applications
Roberto Vitillo

#Distributed_Systems
#network
🧠 یاد گرفتن ساخت سیستمهای توزیعشده سخته، مخصوصاً وقتی در مقیاس بزرگ باشن. مشکل این نیست که اطلاعات کافی وجود نداره. میتونی مقالههای آکادمیک، بلاگهای مهندسی و حتی کتابهایی درباره این موضوع پیدا کنی. مسئله اینه که اطلاعات موجود همهجا پخش شدهاند، و اگر بخوای آنها رو روی طیفی از تئوری تا عمل قرار بدی، میبینی مقدار زیادی محتوا در دو سر این طیف وجود داره، اما وسط طیف چندان پر نیست.
📘 برای همین تصمیم گرفتم کتابی بنویسم که کانسپتهای اصلی تئوری و عملی سیستمهای توزیعشده رو کنار هم بیاره تا لازم نباشه ساعتها وقت بذاری و خودت نقطهها رو به هم وصل کنی. این کتاب تو رو از مبانی سیستمهای توزیعشده در مقیاس بزرگ عبور میده؛ با جزئیات کافی و رفرنسهای بیرونی مناسب برای اینکه هرجا خواستی عمیقتر بشی. این همان راهنماییه که وقتی خودم تازه شروع کرده بودم، آرزو داشتم وجود داشت؛ بر اساس تجربهام در ساخت سیستمهای توزیعشده بزرگی که تا میلیونها Request در ثانیه و میلیاردها دستگاه اسکیل میشن.
👨💻 اگر دولوپری هستی که روی Backend اپلیکیشنهای وب یا موبایل کار میکنی، یا دوست داری وارد این مسیر بشی، این کتاب برای توئه. وقتی اپلیکیشنهای توزیعشده میسازی، باید با Network Stack، مدلهای Data Consistency، پترنهای مقیاسپذیری و Reliability، Best Practiceهای Observability و کلی موضوع دیگه آشنا باشی. البته میتونی بدون دانستن خیلی از اینها هم اپلیکیشن بسازی، اما در نهایت ساعتها برای دیباگ و بازطراحی معماری وقت میذاری و درسهای سختی میگیری که میتونستی خیلی سریعتر و با دردسر کمتر یادشون بگیری.
📌 با این حال، اگر چندین سال تجربه طراحی و ساخت اپلیکیشنهای Highly Available و Fault-Tolerant داری که تا میلیونها کاربر اسکیل میشن، شاید این کتاب برای تو نباشه. بهعنوان یک متخصص، احتمالاً دنبال عمق بیشتری هستی تا گستردگی؛ و این کتاب بیشتر روی گستردگی تمرکز داره، چون در غیر این صورت پوشش دادن این حوزه ممکن نبود.
🔄 ویرایش دوم، بازنویسی کامل ویرایش قبلیه. هر صفحه از ویرایش اول بازبینی شده و هرجا لازم بوده دوباره کار شده؛ بهعلاوه، موضوعهای جدیدی هم برای اولین بار پوشش داده شدهاند.
💬 نظرها
💭 «این کتاب تئوری و Practice دنیای واقعی رو کنار هم میاره و توضیح میده اپلیکیشنهای مقیاس بزرگ چطور از پایه ساخته میشن. با اینکه با بخشی از محتوا آشنا بودم، این اولین باره که میبینم همه چیز اینقدر خوب در قالب یک روایت منسجم کنار هم قرار گرفته. اگر قرار باشه فقط یک کتاب درباره این موضوع بخونی، از همین شروع کن!»
—مائورو دوگلیو، Software Engineer در Microsoft
💭 «راهنمای فوقالعادهایه! یک موضوع پیچیده رو قابلفهم میکنه. کانسپتهای کلیدی رو با تصویرسازیهای عالی و کامل توضیح میده. نشون میده چطور اپلیکیشنهای توزیعشده رو طراحی، ساخت و اجرا کنیم. موضوعهای مهمی مثل Communication، Security، Coordination، Time، Consistency، Transactionها، مقیاسپذیری، Partitioning، Caching، Failureها، Circuit Breakerها، Load Balancing، Operations و Monitoring رو پوشش میده. این کتاب همه اینها رو کنار هم میذاره و یک مرجع معماری محکم ارائه میده.»
—داگ وارن، Technical Editor در Reedsy
💭 «این کتاب جذاب، کانسپتهای آکادمیک رو با تجربه واقعی ساخت سیستمهای توزیعشده در مقیاس بزرگ ترکیب میکنه. چنین سیستمهایی چالشهای زیادی دارن و کتاب تو رو با بعضی از مهمترینها روبهرو میکنه: Load Balancing بین Compute Nodeها، دسترسی و Synchronization داده توزیعشده، مکانیزمهای Resiliency و خیلی چیزهای دیگه.»
—وامیس ژاگیکا، Senior Engineering Manager در Vonage
📖 فهرست مطالب
فصل ۱. مقدمه
بخش ۱. Communication
فصل ۲. لینکهای قابلاعتماد
فصل ۳. لینکهای امن
فصل ۴. Discovery
فصل ۵. APIها
بخش ۲. Coordination
فصل ۶. مدلهای سیستم
فصل ۷. تشخیص Failure
فصل ۸. زمان
فصل ۹. Leader Election
فصل ۱۰. Replication
فصل ۱۱. پرهیز از Coordination
فصل ۱۲. Transactionها
فصل ۱۳. Transactionهای ناهمگام
بخش ۳. Scalability
فصل ۱۴. HTTP Caching
فصل ۱۵. Content Delivery Networkها
فصل ۱۶. Partitioning
فصل ۱۷. ذخیرهسازی فایل
فصل ۱۸. Network Load Balancing
فصل ۱۹. ذخیرهسازی داده
فصل ۲۰. Caching
فصل ۲۱. Microserviceها
فصل ۲۲. Control Planeها و Data Planeها
فصل ۲۳. Messaging
بخش ۴. Resiliency
فصل ۲۴. علتهای رایج Failure
فصل ۲۵. Redundancy
فصل ۲۶. Fault Isolation
فصل ۲۷. Resiliency در Downstream
فصل ۲۸. Resiliency در Upstream
بخش ۵. Maintainability
فصل ۲۹. Testing
فصل ۳۰. Continuous Delivery و Deployment
فصل ۳۱. Monitoring
فصل ۳۲. Observability
فصل ۳۳. Manageability
فصل ۳۴. حرفهای پایانی
👤 درباره نویسنده
👨💻 روبرتو ویتیلو مهندس نرمافزار و نویسندهایه با بیش از یک دهه تجربه در ساخت و رهبری سیستمهای نرمافزاری بزرگمقیاس. کار او حوزههایی مثل مهندسی نرمافزار، رهبری فنی و مدیریت مهندسی رو پوشش میده و همین بهش یک نگاه عملی درباره رفتار اپلیکیشنهای توزیعشده در محیطهای واقعی پروداکشن داده.
📚 او نویسنده کتاب Understanding Distributed Systems است؛ راهنمایی توسعهدهندهمحور که کانسپتهای اصلی تئوری و عملی پشت اپلیکیشنهای توزیعشده بزرگ رو کنار هم میاره. این کتاب برای مهندسهایی نوشته شده که میخوان ایدههای ضروری پشت Communication، Coordination، Scalability، Resilience و Operations رو بفهمن، بدون اینکه در توضیحهای بیش از حد آکادمیک گم بشن.
🛠️ سبک نوشتن ویتیلو از پیشزمینهاش بهعنوان یک مهندس دستبهکار میاد: روشن، عملی و متمرکز روی Trade-offهایی که دولوپرها هنگام طراحی سیستمهایی با نیاز به اسکیل شدن، تحمل Failure و قابلنگهداری ماندن در طول زمان باهاشون روبهرو میشن.
Learning to build distributed systems is hard, especially if they are large scale. It's not that there is a lack of information out there. You can find academic papers, engineering blogs, and even books on the subject. The problem is that the available information is spread out all over the place, and if you were to put it on a spectrum from theory to practice, you would find a lot of material at the two ends but not much in the middle.
That is why I decided to write a book that brings together the core theoretical and practical concepts of distributed systems so that you don't have to spend hours connecting the dots. This book will guide you through the fundamentals of large-scale distributed systems, with just enough details and external references to dive deeper. This is the guide I wished existed when I first started out, based on my experience building large distributed systems that scale to millions of requests per second and billions of devices.
If you are a developer working on the backend of web or mobile applications (or would like to be!), this book is for you. When building distributed applications, you need to be familiar with the network stack, data consistency models, scalability and reliability patterns, observability best practices, and much more. Although you can build applications without knowing much of that, you will end up spending hours debugging and re-architecting them, learning hard lessons that you could have acquired in a much faster and less painful way.
However, if you have several years of experience designing and building highly available and fault-tolerant applications that scale to millions of users, this book might not be for you. As an expert, you are likely looking for depth rather than breadth, and this book focuses more on the latter since it would be impossible to cover the field otherwise.
The second edition is a complete rewrite of the previous edition. Every page of the first edition has been reviewed and where appropriate reworked, with new topics covered for the first time.
"The book brings together theory and real world practice by discussing how large scale applications are built from the ground up. Although I was familiar with some of the content, it's the first time I have seen it woven so well together into a coherent story. If you have to read one book about the topic, start with this one!"
—Mauro Doglio, Software Engineer at Microsoft
"Fantastic guidebook! Makes sense of a complex subject. Thoroughly explains key concepts with great illustrations. Shows how to design, build, and operate distributed applications. It addresses important topics such as communication, security, coordination, time, consistency, transactions, scalability, partitioning, caching, failures, circuit breakers, load balancing, operations, and monitoring. This book puts it all together, and provides a solid architectural reference source."
—Doug Warren, Technical Editor at Reedsy
"This interesting book mixes together academic concepts with real life experience in building distributed systems at scale. Such systems present many challenges, and the book will have you tackle some of the most important ones: load balancing compute nodes, distributed data access and synchronization, resiliency mechanisms, and many more."
—Vamis Xhagjika, Senior Engineering Manager at Vonage
Table of Contents
Chapter 1. Introduction
Part I. Communication
Chapter 2. Reliable Links
Chapter 3. Secure Links
Chapter 4. Discovery
Chapter 5. APIs
Part II. Coordination
Chapter 6. System Models
Chapter 7. Failure Detection
Chapter 8. Time
Chapter 9. Leader Election
Chapter 10. Replication
Chapter 11. Coordination Avoidance
Chapter 12. Transactions
Chapter 13. Asynchronous Transactions
Part III. Scalability
Chapter 14. HTTP Caching
Chapter 15. Content Delivery Networks
Chapter 16. Partitioning
Chapter 17. File Storage
Chapter 18. Network Load Balancing
Chapter 19. Data Storage
Chapter 20. Caching
Chapter 21. Microservices
Chapter 22. Control Planes and Data Planes
Chapter 23. Messaging
Part IV. Resiliency
Chapter 24. Common Failure Causes
Chapter 25. Redundancy
Chapter 26. Fault Isolation
Chapter 27. Downstream Resiliency
Chapter 28. Upstream Resiliency
Part V. Maintainability
Chapter 29. Testing
Chapter 30. Continuous Delivery and Deployment
Chapter 31. Monitoring
Chapter 32. Observability
Chapter 33. Manageability
Chapter 34. Final Words
About the Author
Roberto Vitillo is a software engineer and author with more than a decade of experience building and leading large-scale software systems. His work spans software engineering, technical leadership, and engineering management, giving him a practical view of how distributed applications behave in real production environments.
He is the author of Understanding Distributed Systems, a developer-focused guide that brings together the core theoretical and practical concepts behind large distributed applications. The book is written for engineers who want to understand the essential ideas behind communication, coordination, scalability, resilience, and operations without getting lost in overly academic explanations.
Vitillo’s writing style reflects his background as a hands-on engineer: clear, practical, and focused on the trade-offs developers face when designing systems that need to scale, tolerate failures, and remain maintainable over time.









