Pro Apache Hadoop, Second Edition

7h 26m
Jason Venner, Madhu Siddalingaiah, Sameer Wadkar
Apress
2014

Pro Apache Hadoop, Second Edition brings you up to speed on Hadoop – the framework of big data. Revised to cover Hadoop 2.0, the book covers the very latest developments such as YARN (aka MapReduce 2.0), new HDFS high-availability features, and increased scalability in the form of HDFS Federations. All the old content has been revised too, giving the latest on the ins and outs of MapReduce, cluster design, the Hadoop Distributed File System, and more.

This book covers everything you need to build your first Hadoop cluster and begin analyzing and deriving value from your business and scientific data. Learn to solve big-data problems the MapReduce way, by breaking a big problem into chunks and creating small-scale solutions that can be flung across thousands upon thousands of nodes to analyze large data volumes in a short amount of wall-clock time. Learn how to let Hadoop take care of distributing and parallelizing your software—you just focus on the code; Hadoop takes care of the rest.

Covers all that is new in Hadoop 2.0
Written by a professional involved in Hadoop since day one
Takes you quickly to the seasoned pro level on the hottest cloud-computing framework

What you’ll learn

Build a resilient and scalable Hadoop compute cluster.
Analyze large volumes of data in amazingly short time.
Optimize Hadoop tasks like a seasoned professional.
Implement bulletproof patterns that are proven successful.
Scale out using the new HDFS Federations feature set.
Chunk large problems into highly-parallel, MapReduce modules

Who this book is for

This book is aimed at I.T. professionals investigating Hadoop and implementing it in their organizations. Existing Hadoop users will deepen their toolkits and come up to speed on what’s new Hadoop 2.0. New Hadoop users will quickly move to the seasoned professional level in their use of the toolset.

About the Authors

Jason Venner has more than 20 years of software engineering, managing, designing, and coding experience. He has been a vice president, director, and consultant. Currently, his interests and expertise are in Java, Hadoop, cloud computing, and more.

Sameer Wadkar has over 15 years of experience in Software Development, and over four years of active development experience in Big Data. He has applied Big Data methods in both the Public and Private Sector. He has also applied Big Data and Data Science methods in the Healthcare and Finance Industries. Sameer has a Bachelor in Electrical Engineering from Mumbai University, and a Masters in Applied Mathematics from Johns Hopkins University.

Madhu Siddalingaiah is a technology consultant with 25 years experience in a variety business domains such as aerospace, health care, financial, energy, defense, and scientific research. More recently, Madhu has delivered several high profile big data systems and solutions. Madhu earned his physics degree from the University of MD. Outside of his profession, Madhu is a private helicopter pilot, enjoys travel and a constantly growing list of hobbies.

In this Book

Motivation for Big Data
Hadoop Concepts
Getting Started with the Hadoop Framework
Hadoop Administration
Basics of MapReduce Development
Advanced MapReduce Development
Hadoop Input/Output
Testing Hadoop Programs
Monitoring Hadoop
Data Warehousing Using Hadoop
Data Processing Using Pig
HCatalog and Hadoop in the Enterprise
Log Analysis Using Hadoop
Building Real-Time Systems Using HBase
Data Science with Hadoop
Hadoop in the Cloud
Building a YARN Application

FREE ACCESS

Channel Apache Hadoop

(1)

Book Cloud Native Microservices Cookbook: Master the art of microservices in the cloud with over 100 practical recipes

Course GCP Data Engineer Pro: Creating a Pipeline of Services

(1)

Get Started

Sharpen your skills. Upgrade your career. Find the right learning path for you, based on your role and skills. Take part in hands-on practice, study for a certification, and much more - all personalized for you.

*Not included: Compliance, Leadership Development Program content, and Engineering books

Your content + our content + our platform = a path to learning success

Using our learning experience platform, Percipio, your learners can engage in custom learning paths that can feature curated content from all sources.

Learn More

Aspire to something bigger

Aspire Journeys are guided learning paths that set you in motion for career success.

Browse Aspire Journeys

Explore a world of live learning with Global Knowledge

Choose from convenient delivery formats to get the training you and your team need - where, when and how you want it.

Browse Live Learning

IT Skills and Salary Report

ESG Impact Report

Pro Apache Hadoop, Second Edition

In this Book

YOU MIGHT ALSO LIKE