0
نام کتاب
Serverless ETL and Analytics with AWS Glue

Design scalable data lakes, optimize ETL pipelines, and accelerate analytics on AWS

Noritaka Sekiyama, Albert Quiroga, Tomohiro Tanaka, Subramanya Vajiraya, Akira Ajisaka, Ishan Gaur

Print Length532 Pages
PublisherPackt
Edition2
LanguageEnglish
Year2026
ISBN9781835464847
839
A7196
انتخاب نوع چاپ:
جلد سخت
1,561,000ت
0
جلد نرم
1,691,000ت(2 جلدی)
0
طلق پاپکو و فنر
1,711,000ت(2 جلدی)
0
مجموع:
0تومان
کیفیت متن:اورجینال انتشارات
قطع:B5
رنگ صفحات:دارای متن و کادر رنگی
پشتیبانی در روزهای تعطیل!
ارسال به سراسر کشور

#Serverless

#AWS

#Analytics

#Glue

#ETL

#CI/CD

#CDK

توضیحات

Use AWS Glue to integrate growing data sources with serverless ETL, building secure, observable pipelines that support reliable analytics while managing performance and cost across a governed AWS data platform as workloads grow


Key Features

  • Use runnable code, console walkthroughs, and downloadable examples for core AWS Glue workflows
  • Apply DataOps practices with AWS CDK, Docker, and CI/CD in real-world scenarios
  • Learn from six data specialists with AWS, Spark, Apache Iceberg, and data lake expertise


Book Description

Whether you build data pipelines, design cloud architectures, or deliver analytics on AWS, bringing data together is only part of the challenge. You must also keep this data clean, trustworthy, and available while controlling costs. AWS Glue offers serverless data integration, but using it effectively requires decisions about storage, metadata, security, orchestration, monitoring, and performance.


This book guides you from modern data management and core AWS Glue features through ingestion from files, streams, SaaS applications, and JDBC sources, preparation, storage layout, metadata, security, sharing, and pipeline operations. Console walkthroughs and runnable examples show how to manage schemas and lineage in AWS Glue Data Catalog, apply AWS Lake Formation access controls, monitor workloads, tune Spark jobs, troubleshoot failures, and manage development with AWS CDK, Docker, and CI/CD. You will also examine analytics, machine learning and generative AI integrations, real-world data lake scenarios, and cost optimization. Learn how Apache Iceberg, Apache Hudi, and Delta Lake add transactions, schema evolution, and efficient data management to data lakes.


By the end, you will be able to design, build, operate, and continuously improve a serverless data platform that fits your organization's scale, structure, and priorities.


What you will learn

  • Design scalable serverless ETL pipelines with AWS Glue
  • Ingest data from files, streams, SaaS, and JDBC sources
  • Optimize file formats, partitions, compression, and layouts
  • Manage schemas, partitions, and lineage in AWS Glue Data Catalog
  • Secure data with access control, encryption, and auditing
  • Automate testing and multi account CI/CD using AWS CDK and Docker
  • Monitor, tune, and troubleshoot AWS Glue and Spark workloads
  • Apply Apache Iceberg, Hudi, and Delta Lake to data lakes with AWS Glue


Who this book is for

This book is for data engineers, ETL developers, cloud architects, and analytics professionals who build or operate data platforms on AWS. It suits readers working on serverless data lakes, Spark ETL, governance, data sharing, reliability, or cost control. It is especially useful if you aim to improve pipeline reliability, governance, or cost visibility as workloads grow. Basic familiarity with the AWS Management Console, Amazon S3, and IAM is recommended. Experience with Python, SQL, or Apache Spark will help with the code examples, and an AWS account is useful for following the walkthroughs.


Table of Contents

Part I: Introduction, Concepts and Basics of AWS Glue

Chapter 1: Data Management – Introduction and Concepts

Chapter 2: Introduction to Important AWS Glue Features

Chapter 3: Data Ingestion

Part II: Data Preparation, Management and Security

Chapter 4: Data Preparation

Chapter 5: Data Layouts

Chapter 6: Data Management

Chapter 7: Implementing Resilient Metadata Management

Chapter 8: Data Security

Chapter 9: Data Sharing

Chapter 10: Data Pipeline Management

Part III: Tuning, Monitoring, and Real-World Scenarios

Chapter 11: Monitoring

Chapter 12: Tuning, Debugging, and Troubleshooting

Chapter 13: Data Analysis

Chapter 14: Machine Learning and Generative AI Integration

Chapter 15: Architecting Data Lakes for Real-World Scenarios and Edge Cases

Chapter 16: Managing End-to-End Development Lifecycle

Chapter 17: Open Table Formats

Chapter 18: Cost Optimization


About the Authors

Noritaka Sekiyama is an experienced big data engineer working at a data and AI company. He is responsible for building scalable data platforms with unified governance in the cloud. He is passionate about software engineering, cloud computing, big data technologies, distributed systems, data platforms, system monitoring, and automation.


Albert Quiroga is a Senior Solutions Architect at Amazon, where he creates solutions and architectural designs for one of the largest data lakes in the world. Prior to that, he spent four years working at AWS, where he specialized in big data technologies such as Amazon EMR, Amazon Athena, AWS Glue, and Amazon SageMaker. His 11 years of experience in the industry have empowered him to work with several Fortune 500 companies to overcome large-scale data and analytics challenges, and he has helped launch and develop features for several AWS services.


Tomohiro Tanaka is a big data specialist with deep, hands-on expertise in data infrastructure. His expertise covers large-scale migrations, performance tuning, and production troubleshooting, with a focus on Apache Spark and Apache Iceberg. He contributes to the Apache Iceberg open-source project and speaks at community events and conferences to help teams adopt Apache Iceberg in practice.

Subramanya Vajiraya is a Senior Cloud Engineer at AWS Sydney specializing in AWS Glue. He obtained his Bachelor of Engineering degree in Information Science & Engineering from NMAM Institute of Technology, Nitte, KA, India, in 2015 and his Master of Information Technology degree in Internetworking from the University of New South Wales, Sydney, Australia, in 2017. He is passionate about helping customers solve challenging technical issues related to their ETL workloads and implement scalable data integration and analytics pipelines on AWS.


Akira Ajisaka is a software engineer with more than 10 years of engineering experience in big data. He enjoys troubleshooting and contributing to OSS.


Ishan Gaur has more than 17 years of IT experience in software development, data engineering, and cloud architecture, building distributed systems and highly scalable data processing pipelines using Apache Spark, Scala, and various AWS data services, such as AWS Glue, Amazon SageMaker Unified Studio, and Amazon EMR. He currently works at AWS as a Principal Cloud Engineer, where he is focused on AI/ML operations and proactive cloud optimization. He works with AWS enterprise customers to design resilient data pipelines, automate incident response, troubleshoot large-scale distributed data platforms, and adopt GenAI-powered services and operational tools. He is passionate about turning reactive support patterns into proactive, self-healing architectures.

دیدگاه خود را بنویسید
نظرات کاربران (0 دیدگاه)
نظری وجود ندارد.
کتاب های مشابه
AWS
1,221
AWS Certified Cloud Practitioner Study Guide
971,000 تومان
AWS
1,009
Serverless Architectures on AWS
826,000 تومان
AWS
1,181
Hands-On AWS Penetration Testing with Kali Linux
1,436,000 تومان
Artificial intelligence
1,047
Generative AI on AWS
947,000 تومان
AWS
1,005
Building Multi-Tenant SaaS Architectures
1,332,000 تومان
AWS
960
AWS Certified SysOps Administrator - Associate (SOA-C02) Exam Cram
1,008,000 تومان
AWS
2,699
Accelerating DevSecOps on AWS
1,768,000 تومان
AWS
1,148
Programming AWS Lambda
872,000 تومان
AWS
1,109
Using Amazon Bedrock
1,305,000 تومان
Machine Learning
1,628
Automated Machine Learning on AWS
1,271,000 تومان
قیمت
منصفانه
ارسال به
سراسر کشور
تضمین
کیفیت
پشتیبانی در
روزهای تعطیل
خرید امن
و آسان
آرشیو بزرگ
کتاب‌های تخصصی
هـر روز با بهتــرین و جــدیــدتـرین
کتاب های روز دنیا با ما همراه باشید
آدرس
پشتیبانی
مدیریت
ساعات پاسخگویی
درباره اسکای بوک
دسترسی های سریع
  • راهنمای خرید
  • راهنمای ارسال
  • سوالات متداول
  • قوانین و مقررات
  • وبلاگ
  • درباره ما