A structured guide to NVIDIA GPU infrastructure, from CUDA to production operations
Vivian Aranha

#NVIDIA
#GPU
#GPU
#vGPU
#CUDA
#MLOps
#DOCA
#MIG
#NGC
#DPU
#RDMA
#ONNX
🟢 راهنمای جامع زیرساخت NVIDIA GPU برای هوش مصنوعی
⚙️ این کتاب اکوسیستم پیچیده NVIDIA GPU را در قالب مسیری منظم و یکپارچه توضیح میده. با مقایسه فناوریها و بررسی ارتباط میان لایههای مختلف پلتفرم، یاد میگیری انتخابهای زیرساختی و Trade-offهای فنی را با دید دقیقتری ارزیابی کنی.
✨ ویژگیهای کلیدی
🧠 توضیح میده GPUها، CUDA، Networking، Storage و DPUها چطور از AI Workloadها پشتیبانی میکنن.
🧩 نقش ابزارها و فناوریهایی مثل MIG، vGPU، DCGM، Kubernetes، Slurm، NGC و Triton را در زیرساخت GPU مشخص میکنه.
🔄 اجزای Infrastructure را در سراسر چرخه توسعه، آموزش، استقرار و اجرای مدلهای AI به هم متصل میکنه.
📘 توضیح کتاب
🏗️ زیرساخت NVIDIA GPU فقط شامل Hardware نیست و مجموعهای از System Software، Networking، Storage، Orchestration، MLOps و فناوریهای Inference را در بر میگیره. شناخت ارتباط این اجزا، محدوده مسئولیت آنها و نقاط همپوشانیشان به یک مسیر آموزشی روشن و ساختاریافته نیاز داره.
⚡ کتاب این مسیر را با ارائه روایتی منسجم از NVIDIA GPU Infrastructure Stack فراهم میکنه. مطالب با Accelerated Computing، CUDA و اکوسیستم نرمافزاری NVIDIA شروع میشن و بعد GPUهای Data Center براساس ویژگیهای Workloadهای مختلف مقایسه میشن.
🖥️ در ادامه، فناوریهای MIG و vGPU برای Resource Sharing و DCGM برای Monitoring بررسی میشن. کتاب همچنین نقش Kubernetes و Slurm را در Scheduling منابع GPU توضیح میده.
🌐 لایه Networking و Storage نیز با موضوعاتی مثل Ethernet، InfiniBand، RDMA، GPUDirect Storage، BlueField DPUها و DOCA پوشش داده میشه. این بخش کمک میکنه مسیر انتقال داده و تأثیر آن بر کارایی GPU Clusterها را بهتر درک کنی.
🔄 فصلهای بعدی Infrastructure را به AI Lifecycle متصل میکنن. Airflow، MLflow و Kubeflow برای MLOps، پلتفرم NGC برای Software Delivery و فناوریهای ONNX، TensorRT و Triton برای Inference بررسی میشن.
📊 کتاب علاوه بر اجزای Kubernetes، فناوریهای Monitoring، ملاحظات Scaling و مفاهیم Diagnostic موردنیاز برای پشتیبانی از Production GPU Clusterها را پوشش میده. فناوریها بهعنوان محصولاتی جدا از هم معرفی نمیشن، بلکه ارتباط و مرز مسئولیت آنها مشخص میشه.
🎯 بعد از مطالعه کتاب، میتونی انتخابهای مختلف پلتفرم را مقایسه کنی، محدوده عملکرد هر Component را بشناسی، درباره Trade-offها دقیقتر صحبت کنی و یک Mental Model ماندگار از زیرساخت NVIDIA GPU به دست بیاری.
🎯 چیزهایی که یاد میگیری
🧠 یاد میگیری تفاوت میان AI، Machine Learning و Deep Learning را تشخیص بدی.
⚡ درک میکنی چرا GPUها باعث شتابگرفتن AI Workloadهای مدرن میشن.
🖥️ میتونی GPUهای NVIDIA را با نیازهای Training و Inference تطبیق بدی.
🧩 یاد میگیری برای سناریوهای رایج Resource Sharing میان MIG و vGPU انتخاب کنی.
🌐 میتونی Ethernet و InfiniBand را برای Distributed AI Workloadها مقایسه کنی.
🔄 یاد میگیری ابزارهای MLOps را به مرحله مناسب AI Lifecycle متصل کنی.
🚀 تفاوت نقش ONNX، TensorRT و Triton را در Inference Workflowها درک میکنی.
🔍 میتونی مشکلات GPU Cluster را در لایههای مختلف پلتفرم ردیابی و بررسی کنی.
👤 این کتاب برای چه کسانیه؟
🖥️ این کتاب برای System Administratorها، متخصصان Cloud و DevOps، تیمهای Data Center و Networking، Solution Architectها، مدیران فنی و متخصصان Presales نوشته شده که به درکی روشن از زیرساخت NVIDIA GPU نیاز دارن.
🤖 مطالب برای افراد مبتدی و متخصصانی هم مناسبه که میخوان وارد نقشهای AI Infrastructure بشن، فناوریهای GPU Platform را ارزیابی کنن یا میان تیمهای Compute، Networking، MLOps و Operations همکاری داشته باشن.
📌 آشنایی پایه با مفاهیم IT، Cloud یا Data Center مفیده؛ اما برای مطالعه کتاب به تجربه برنامهنویسی، Data Science یا کار قبلی با GPU نیازی نیست.
📖 فهرست مطالب
فصل ۱. شناخت AI Workloadها و Accelerated Computing
فصل ۲. CUDA و NVIDIA AI Software Stack
فصل ۳. معماری NVIDIA GPU و انتخاب پلتفرم
فصل ۴. اشتراکگذاری، Monitoring و Scheduling منابع GPU
فصل ۵. ساخت مسیرهای پرظرفیت Storage و Network
فصل ۶. زیرساخت Virtualized GPU و DPU Offload
فصل ۷. مدیریت AI Lifecycle با MLOps و Inference
فصل ۸. ارائه و مدیریت NVIDIA AI Platformها
فصل ۹. آمادگی برای آزمون NCA-AIIO و گام بعدی
👤 درباره نویسنده
🤖 ویویان آرانها مدرس AI، مدیر فناوری و بنیانگذار School of AI است و بیش از ۲۰ سال تجربه فعالیت در صنعت داره.
🎓 او مدرک کارشناسی Information Technology را در سال ۲۰۰۴ و مدرک کارشناسی ارشد Computer Science را در سال ۲۰۰۶ دریافت کرده است.
💻 مسیر حرفهای ویویان آرانها شامل Web Technology، توسعه Mobile Application برای iOS و Android، راهکارهای Blockchain و سیستمها و Applicationهای AI میشه.
🏢 او با سازمانهای Fortune 500، از جمله The Washington Post، Delta Air Lines و IBM همکاری کرده است.
📚 ویویان از سال ۲۰۰۹ بهعنوان Instructor فعالیت میکنه، متخصصان بسیاری را در سراسر جهان آموزش داده و اکنون AI را در سطح بینالمللی تدریس میکنه.
🌍 دورههای آموزشی او بیش از ۲.۵ میلیون ثبتنام داشتهاند و بیشتر از ۵۰۰ هزار دانشجو از طریق School of AI، Udemy، Skool و Maven در دورههای او آموزش دیدهاند.
Decode the NVIDIA GPU ecosystem in one structured guide. Compare technologies, understand how platform layers interact, and build the judgment to evaluate infrastructure choices and trade-offs.
NVIDIA GPU infrastructure spans hardware, system software, networking, storage, orchestration, MLOps, and inference. Understanding how these components fit together, where their responsibilities overlap, and which distinctions matter requires a clear, structured path.
This book provides that path through one coherent narrative of the NVIDIA GPU infrastructure stack. It covers accelerated computing, CUDA, and the NVIDIA software ecosystem before comparing data center GPUs against workload characteristics. You will examine MIG, vGPU, and DCGM for resource sharing and monitoring; Kubernetes and Slurm for GPU scheduling; and the networking and storage layer, including Ethernet, InfiniBand, RDMA, GPUDirect Storage, BlueField DPUs, and DOCA.
Later chapters connect infrastructure to the AI lifecycle through Airflow, MLflow, and Kubeflow for MLOps, NGC for software delivery, and ONNX, TensorRT, and Triton for inference. You will also explore the Kubernetes components, monitoring technologies, scaling considerations, and diagnostic concepts that support production GPU clusters.
By connecting these technologies instead of presenting them as isolated products, the book helps you compare platform choices, understand component boundaries, discuss trade-offs, and develop a durable mental model of NVIDIA GPU infrastructure.
This book is for system administrators, cloud and DevOps professionals, data center and networking teams, solution architects, technical managers, presales professionals, and beginners who need a clear understanding of NVIDIA GPU infrastructure. It is especially relevant to professionals moving into AI infrastructure roles, evaluating GPU platform technologies, or collaborating across compute, networking, MLOps, and operations teams. Basic familiarity with IT, cloud, or data center concepts is helpful; programming, data science, and previous GPU experience are not required.
About the Author
Vivian Aranha is an AI educator, technology leader, and founder of School of AI, with over 20 years of industry experience. He earned a Bachelor's degree in Information Technology in 2004 and a Master's degree in Computer Science in 2006. His career spans web technologies, mobile app development for iOS and Android, blockchain solutions, and AI systems and applications. Vivian has worked with Fortune 500 organizations, including The Washington Post, Delta Air Lines, and IBM. An instructor since 2009, he has trained professionals worldwide and now teaches AI globally. His courses have attracted over 2.5 million enrollments, with more than 500,000 students learning through School of AI, Udemy, Skool, and Maven.









