Carlos Fernandez-Granda

#Probability
#Statistics
#Data_Science
📘 این راهنمای مستقل، دو ستون اصلی علم داده یعنی نظریه احتمال و آمار رو در کنار هم معرفی میکنه تا ارتباط بین تکنیکهای آماری و مفاهیم احتمالاتیای که پایه اونها هستن، روشنتر بشه.
🧠 کتاب موضوعهایی مثل متغیرهای تصادفی، مدلهای پارامتری و ناپارامتری، همبستگی، برآورد پارامترهای جامعه آماری، آزمون فرض، تحلیل مؤلفههای اصلی و روشهای خطی و غیرخطی برای رگرسیون و طبقهبندی رو پوشش میده.
🌍 مثالهای سراسر کتاب از دیتاستهای واقعی استفاده میکنن تا مفاهیم رو در عمل نشون بدن و خواننده رو با چالشهای بنیادی علم داده مثل بیشبرازش، نفرین ابعاد و استنباط علّی روبهرو کنن.
🐍 کدهای Python مربوط به مثالها در سایت کتاب در دسترسن و میتونی همراه اونها از ویدئوها، اسلایدها و پاسخ تمرینها هم استفاده کنی.
🎓 توضیحهای قابلفهم کتاب باعث شده این اثر برای دانشجوهای کارشناسی و تحصیلات تکمیلی، متخصصهای علم داده و هر کسی که به مفاهیم نظری پشت روشهای علم داده علاقه داره، منبعی مناسب باشه.
📖 فهرست مطالب
فصل ۱. احتمال
فصل ۲. متغیرهای گسسته
فصل ۳. متغیرهای پیوسته
فصل ۴. چند متغیر گسسته
فصل ۵. چند متغیر پیوسته
فصل ۶. متغیرهای گسسته و پیوسته
فصل ۷. میانگینگیری
فصل ۸. همبستگی
فصل ۹. برآورد پارامترهای جامعه آماری
فصل ۱۰. آزمون فرض
فصل ۱۱. تحلیل مؤلفههای اصلی و مدلهای کمرتبه
فصل ۱۲. رگرسیون و طبقهبندی
💬 نظرها
💭 «کتاب Probability and Statistics for Data Science نوشته فرناندز-گراندا، مباحث بنیادی موردنیاز تمام کسانی رو که میخوان وارد علم داده بشن، بهشکلی جامع و در عین حال قابلفهم ارائه میده؛ چه در محیط دانشگاهی فعالیت کنن، چه در صنعت یا حوزههای دیگه.
زبان کتاب روشن و دقیق است و ساختار اون یکی از منظمترین ارائههایی محسوب میشه که تا حالا از این مطالب دیدهام. مثالهای شفاف و تمرینهای مفید باعث شده این کتاب شایستگی تبدیل شدن به منبع اصلی این موضوعها برای دانشجوهای کارشناسی و تحصیلات تکمیلی در این رشته فنی و سریعالتغییر رو داشته باشه. مدرسها باید حتماً به این کتاب توجه کنن.»
—آرتور اسپیرلینگ، Princeton University
💭 «اگر به ریاضی علاقه داری و میخوای مبانی علم داده رو یکجا و کامل یاد بگیری، این کتاب برای توئه. دامنه گستردهای از موضوعهای ضروری و مدرن رو پوشش میده؛ از روشهای ناپارامتری و استنباط علّی گرفته تا مدلهای متغیر پنهان، رویکردهای بیزی و مقدمهای کامل بر یادگیری ماشین.
همه این موضوعها با تعداد زیادی شکل و مثال مبتنی بر دادههای واقعی توضیح داده شدهاند. این کتاب رو بهشدت پیشنهاد میکنم.»
—دیوید روزنبرگ، دفتر مدیر ارشد فناوری Bloomberg
📗 توضیح کوتاه کتاب
یک مقدمه مستقل و کامل درباره احتمال و آمار برای علم داده که مفاهیم رو با مثالهای مبتنی بر دیتاستهای واقعی آموزش میده.
👤 درباره نویسنده
👨🏫 کارلوس فرناندز-گراندا دانشیار ریاضیات و علم داده در New York University است و از سال ۲۰۱۵ احتمال و آمار رو به دانشجوهای علم داده آموزش میده.
🔬 هدف پژوهشهای او طراحی و تحلیل روشهای علم داده است و تمرکز ویژهای روی یادگیری ماشین، هوش مصنوعی و کاربرد اونها در پزشکی، علوم اقلیمی، زیستشناسی و دیگر حوزههای علمی داره.
This self-contained guide introduces two pillars of data science, probability theory, and statistics, side by side, in order to illuminate the connections between statistical techniques and the probabilistic concepts they are based on. The topics covered in the book include random variables, nonparametric and parametric models, correlation, estimation of population parameters, hypothesis testing, principal component analysis, and both linear and nonlinear methods for regression and classification. Examples throughout the book draw from real-world datasets to demonstrate concepts in practice and confront readers with fundamental challenges in data science, such as overfitting, the curse of dimensionality, and causal inference. Code in Python reproducing these examples is available on the book's website, along with videos, slides, and solutions to exercises. This accessible book is ideal for undergraduate and graduate students, data science practitioners, and others interested in the theoretical concepts underlying data science methods.
Table of Contents
1 Probability
2 Discrete Variables
3 Continuous Variables
4 Multiple Discrete Variables
5 Multiple Continuous Variables
6 Discrete and Continuous Variables
7 Averaging
8 Correlation
9 Estimation of Population Parameters
10 Hypothesis Testing
11 Principal Component Analysis and Low-Rank Models
12 Regression and Classification
‘Fernandez-Granda's Probability and Statistics for Data Science is a comprehensive yet approachable treatment of the fundamentals required of all aspiring Data Scientists-whether they be in academia, industry or elsewhere. The language is clear and precise, and it is one of the best-organized treatments of this material I have ever seen. With lucid examples and helpful exercises, it deserves to be the leading text for these topics among undergraduate and graduate students in this technical, fast-moving discipline. Instructors take note!’ Arthur Spirling, Princeton University
‘If you're mathematically inclined and want to master the foundations of data science in one go, this book is for you. It covers a broad range of essential modern topics - including nonparametric methods, causal inference, latent variable models, Bayesian approaches, and a thorough introduction to machine learning - all illustrated with an abundance of figures and real-world data examples. Highly recommended.’ David Rosenberg, Office of the CTO, Bloomberg
A self-contained introduction to probability and statistics for data science with examples involving real-world datasets.
Carlos Fernandez-Granda is Associate Professor of Mathematics and Data Science at New York University, where he has taught probability and statistics to data science students since 2015. The goal of his research is to design and analyze data science methodology, with a focus on machine learning, artificial intelligence, and their application to medicine, climate science, biology, and other scientific domains.









