Scaling Big Data with Hadoop and Solr - Second Edition
eBook - ePub

Scaling Big Data with Hadoop and Solr - Second Edition

  1. 166 pages
  2. English
  3. ePUB (mobile friendly)
  4. Available on iOS & Android
eBook - ePub

Scaling Big Data with Hadoop and Solr - Second Edition

About this book

About This Book

  • Explore different approaches to making Solr work on big data ecosystems besides Apache Hadoop
  • Improve search performance while working with big data
  • A practical guide that covers interesting, real-life use cases for big data search along with sample code

Who This Book Is For

This book is aimed at developers, designers, and architects who would like to build big data enterprise search solutions for their customers or organizations. No prior knowledge of Apache Hadoop and Apache Solr/Lucene technologies is required.

Trusted by 375,005 students

Access to over 1.5 million titles for a fair monthly price.

Study more efficiently using our study tools.

Information

Year
2015
eBook ISBN
9781783553402

Scaling Big Data with Hadoop and Solr Second Edition


Table of Contents

Scaling Big Data with Hadoop and Solr Second Edition
Credits
About the Author
About the Reviewers
www.PacktPub.com
Support files, eBooks, discount offers, and more
Why subscribe?
Free access for Packt account holders
Preface
What this book covers
What you need for this book
Who this book is for
Conventions
Reader feedback
Customer support
Downloading the example code
Errata
Piracy
Questions
1. Processing Big Data Using Hadoop and MapReduce
Apache Hadoop's ecosystem
Core components
Understanding Hadoop's ecosystem
Configuring Apache Hadoop
Prerequisites
Setting up ssh without passphrase
Configuring Hadoop
Running Hadoop
Setting up a Hadoop cluster
Common problems and their solutions
Summary
2. Understanding Apache Solr
Setting up Apache Solr
Prerequisites for setting up Apache Solr
Running Apache Solr on jetty
Running Solr on other J2EE containers
Hello World with Apache Solr!
Understanding Solr administration
Solr navigation
Common problems and solutions
The Apache Solr architecture
Configuring Solr
Understanding the Solr structure
Defining the Solr schema
Solr fields
Dynamic fields in Solr
Copying the fields
Dealing with field types
Additional metadata configuration
Other important elements of the Solr schema
Configuration files of Apache Solr
Working with solr.xml and Solr core
Instance configuration with solrconfig.xml
Understanding the Solr plugin
Other configuration
Loading data in Apache Solr
Extracting request handler – Solr Cell
Understanding data import handlers
Interacting with Solr through SolrJ
Working with rich documents (Apache Tika)
Querying for information in Solr
Summary
3. Enabling Distributed Search using Apache Solr
Understanding a distributed search
Distributed search patterns
Apache Solr and distributed search
Working with SolrCloud
Why ZooKeeper?
The SolrCloud architecture
Building an enterprise distributed search using SolrCloud
Setting up SolrCloud for development
Setting up SolrCloud for production
Adding a document to SolrCloud
Creating shards, collections, and replicas in SolrCloud
Common problems and resolutions
Sharding algorithm and fault tolerance
Document Routing and Sharding
Shard splitting
Load balancing and fault tolerance in SolrCloud
Apache Solr and Big Data – integration with MongoDB
What is NoSQL and how is it related to Big Data?
MongoDB at glance
Installing MongoDB
Creating Solr indexes from MongoDB
Summary
4. Big Data Search Using Hadoop and Its Ecosystem
Understanding NoSQL
Working with the Solr HDFS connector
Big data search using Katta
How Katta works?
Setting up the Katta cluster
Creating Katta indexes
Using Solr 1045 Patch – map-side indexing
Using Solr 1301 Patch – reduce-side indexing
Distributed search using Apache Blur
Setting up Apache Blur with Hadoop
Apache Solr and Cassandra
Working with Cassandra and Solr
Single node configuration
Integrating with multinode Cassandra
Scaling Solr through Storm
Getting along with Apache Storm
Advanced analytics with Solr
Integrating Solr and R
Summary
5. Scaling Search Performance
Understanding the limits
Optimizing search schema
Specifying default search field
Configuring search schema fields
Stop words
Stemming
Index optimization
Limiting indexing buffer size
When to commit changes?
Optimizing index merge
Optimize option for index merging
Optimizing the container
Optimizing concurrent clients
Optimizing Java virtual memory
Optimizing search runtime
Optimizing through search query
Filter queries
Optimizing the Solr cache
The filter cache
The query result cache
The document cache
The field value cache
The lazy field loading
Optimizing Hadoop
Monitoring Solr instance
Using SolrMeter
Summary
A. Use Cases for Big Data Search
E-Commerce websites
Log management for banking
The problem
How can it be tackled?
High-level design
Index

Scaling Big Data with Hadoop and Solr Second Edition

Copyright © 2015 Packt Publishing
All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews.
Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book.
Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information.
First published: August 2013
Second edition: April 2015
Production reference: 1230415
Published by Packt Publishing Ltd.
Livery Place
35 Livery Street
Birmingham B3 2PB, UK.
ISBN 978-1-78355-339-6
www.packtpub.com

Credits

Author
Hrishikesh Vijay Karambelkar
Reviewers
Ramzi Alqrainy
Walt Stoneburner
Ning Sun
Ruben Teijeiro
Commissioning Editor
Kartikey Pandey
Acquisition Editor
Nikhil Chinnari
Reshma Raman
Content Development Editor
Susmita Sabat
Technical Editor
Aman Preet Singh
Copy Editors
Sonia Cheema
Tani Kothari
Project Coordinator
Milton Dsouza
Proofreader
Simran Bhogal
Safis Editing
Indexer
Mariammal Chettiyar
Production Coordinator
Arvindkumar Gupta
Co...

Table of contents

  1. Scaling Big Data with Hadoop and Solr Second Edition

Frequently asked questions

Yes, you can cancel anytime from the Subscription tab in your account settings on the Perlego website. Your subscription will stay active until the end of your current billing period. Learn how to cancel your subscription
No, books cannot be downloaded as external files, such as PDFs, for use outside of Perlego. However, you can download books within the Perlego app for offline reading on mobile or tablet. Learn how to download books offline
Perlego offers two plans: Essential and Complete
  • Essential is ideal for learners and professionals who enjoy exploring a wide range of subjects. Access the Essential Library with 800,000+ trusted titles and best-sellers across business, personal growth, and the humanities. Includes unlimited reading time and Standard Read Aloud voice.
  • Complete: Perfect for advanced learners and researchers needing full, unrestricted access. Unlock 1.5M+ books across hundreds of subjects, including academic and specialized titles. The Complete Plan also includes advanced features like Premium Read Aloud and Research Assistant.
Both plans are available with monthly, semester, or annual billing cycles.
We are an online textbook subscription service, where you can get access to an entire online library for less than the price of a single book per month. With over 1.5 million books across 990+ topics, we’ve got you covered! Learn about our mission
Look out for the read-aloud symbol on your next book to see if you can listen to it. The read-aloud tool reads text aloud for you, highlighting the text as it is being read. You can pause it, speed it up and slow it down. Learn more about Read Aloud
Yes! You can use the Perlego app on both iOS and Android devices to read anytime, anywhere — even offline. Perfect for commutes or when you’re on the go.
Please note we cannot support devices running on iOS 13 and Android 7 or earlier. Learn more about using the app
Yes, you can access Scaling Big Data with Hadoop and Solr - Second Edition by Hrishikesh Vijay Karambelkar in PDF and/or ePUB format, as well as other popular books in Computer Science & Programming in Java. We have over 1.5 million books available in our catalogue for you to explore.