Book

Essential PySpark for Scalable Data Analytics

Essential PySpark for Scalable Data Analytics is an introduction for anyone new to the distributed computing model. You'll learn to unlock the analytics world by building end-to-end data processing pipelines, starting with data ingestion, cleansing, and integration, through to data visualization and building and operationalizing predictive models.

Offered by

Difficulty Level

Intermediate

Completion Time

10h44m

Language

English

About Book

Who Is This Book For?

This book is for practicing data engineers, data scientists, data analysts, and data enthusiasts who are already using data analytics to explore distributed and scalable data analytics. Basic to intermediate knowledge of the disciplines of data engineering, data science, and SQL analytics is expected. General proficiency in using any programming language, especially Python, and working knowledge of performing data analytics using frameworks such as pandas and SQL will help you to get the most out of this book.