MTU Library Catalogue

Syndetics cover image
Image from Syndetics

Big data : principles and best practices of scalable real-time data systems / Nathan Marz, James Warren.

By: Marz, Nathan [author].
Contributor(s): Warren, James (Software engineer) [author].
Material type: materialTypeLabelBookPublisher: Shelter Island, NY : Manning, [2015]Copyright date: ©2015Description: xx, 308 pages : illustrations ; 24 cm.Content type: text Media type: unmediated Carrier type: volumeISBN: 9781617290343; 1617290343.Subject(s): Big data | Database management | Database design | Data miningAdditional physical formats: Electronic Version: Big data : principles and best practices of scalable real-time data systemsDDC classification: 658.4038 Also available in electronic form
Contents:
A new paradigm for big data -- Data model for big data -- Data model for big data : illustration -- Data storage on the batch layer -- Data storage on the batch layer : illustration -- Batch layer -- Batch layer : illustration -- An example batch layer : architecture and algorithms -- An example batch layer : implementation -- Serving layer -- Serving layer : illustration -- Realtime views -- Realtime views : illustration -- Queuing and stream processing -- Queuing and stream processing : illustration -- Micro-batch stream processing -- Micro-batch stream processing : illustration -- Lambda Architecture in depth.
Holdings
Item type Current library Call number Status Notes Barcode
General lending MTU Bishopstown Library Lending 658.4038 (Browse shelf(Opens below)) Available CIT Module INFO8012 - Core reading. 00162819
Total holds: 0

Enhanced descriptions from Syndetics:

Services like social networks, web analytics, and intelligent e-commerce often need to manage data at a scale too big for a traditional database. As scale and demand increase, so does Complexity. Fortunately, scalability and simplicity are not mutually exclusive-- rather than using some trendy technology, a different approach is needed. Big data systems use many machines working in parallel to store and process data, which introduces fundamental challenges unfamiliar to most developers.

Big Data shows how to build these systems using an architecture that takes advantage of clustered hardware along with new tools designed specifically to capture and analyze web-scale data. It describes a scalable, easy to understand approach to big data systems that can be built and run by a small team. Following a realistic example, this book guides readers through the theory of big data systems, how to use them in practice, and how to deploy and operate them once they're built.

AUDIENCE

This book requires no previous exposure to large-scale data analysis or NoSQL tools. Familiarity with traditional databases is helpful.

ABOUT THE TECHNOLOGY

To tackle the challenges of Big Data, a new breed of technologies has emerged. Many of which have been grouped under the term "NoSQL." In some ways these new technologies can be more complex than traditional databases and in other ways, simpler. Using them effectively requires a fundamentally new set of techniques

Includes index.

A new paradigm for big data -- Data model for big data -- Data model for big data : illustration -- Data storage on the batch layer -- Data storage on the batch layer : illustration -- Batch layer -- Batch layer : illustration -- An example batch layer : architecture and algorithms -- An example batch layer : implementation -- Serving layer -- Serving layer : illustration -- Realtime views -- Realtime views : illustration -- Queuing and stream processing -- Queuing and stream processing : illustration -- Micro-batch stream processing -- Micro-batch stream processing : illustration -- Lambda Architecture in depth.

CIT Module INFO8012 - Core reading.

Also available in electronic form

Author notes provided by Syndetics

Nathan Marz is an engineer at Twitter. He was previously Lead Engineer at BackType, a marketing intelligence company that was acquired by Twitter in July of 2011. He is the author of two major open source projects: Storm, a distributed realtime computation system, and Cascalog, a tool for processing data on Hadoop. He is a frequent speaker and writes a blog at nathanmarz.com. James Warren is an analytics architect at Storm8 with a background in big data processing, machine learning and scientific computing.