Scaling Big Data with Hadoop and Solr - Second Edition
eBook - ePub

Scaling Big Data with Hadoop and Solr - Second Edition

Hrishikesh Vijay Karambelkar

Buch teilen
  1. 166 Seiten
  2. English
  3. ePUB (handyfreundlich)
  4. Über iOS und Android verfügbar
eBook - ePub

Scaling Big Data with Hadoop and Solr - Second Edition

Hrishikesh Vijay Karambelkar

Angaben zum Buch
Buchvorschau
Inhaltsverzeichnis
Quellenangaben

Über dieses Buch

About This Book

  • Explore different approaches to making Solr work on big data ecosystems besides Apache Hadoop
  • Improve search performance while working with big data
  • A practical guide that covers interesting, real-life use cases for big data search along with sample code

Who This Book Is For

This book is aimed at developers, designers, and architects who would like to build big data enterprise search solutions for their customers or organizations. No prior knowledge of Apache Hadoop and Apache Solr/Lucene technologies is required.

Häufig gestellte Fragen

Wie kann ich mein Abo kündigen?
Gehe einfach zum Kontobereich in den Einstellungen und klicke auf „Abo kündigen“ – ganz einfach. Nachdem du gekündigt hast, bleibt deine Mitgliedschaft für den verbleibenden Abozeitraum, den du bereits bezahlt hast, aktiv. Mehr Informationen hier.
(Wie) Kann ich Bücher herunterladen?
Derzeit stehen all unsere auf Mobilgeräte reagierenden ePub-Bücher zum Download über die App zur Verfügung. Die meisten unserer PDFs stehen ebenfalls zum Download bereit; wir arbeiten daran, auch die übrigen PDFs zum Download anzubieten, bei denen dies aktuell noch nicht möglich ist. Weitere Informationen hier.
Welcher Unterschied besteht bei den Preisen zwischen den Aboplänen?
Mit beiden Aboplänen erhältst du vollen Zugang zur Bibliothek und allen Funktionen von Perlego. Die einzigen Unterschiede bestehen im Preis und dem Abozeitraum: Mit dem Jahresabo sparst du auf 12 Monate gerechnet im Vergleich zum Monatsabo rund 30 %.
Was ist Perlego?
Wir sind ein Online-Abodienst für Lehrbücher, bei dem du für weniger als den Preis eines einzelnen Buches pro Monat Zugang zu einer ganzen Online-Bibliothek erhältst. Mit über 1 Million Büchern zu über 1.000 verschiedenen Themen haben wir bestimmt alles, was du brauchst! Weitere Informationen hier.
Unterstützt Perlego Text-zu-Sprache?
Achte auf das Symbol zum Vorlesen in deinem nächsten Buch, um zu sehen, ob du es dir auch anhören kannst. Bei diesem Tool wird dir Text laut vorgelesen, wobei der Text beim Vorlesen auch grafisch hervorgehoben wird. Du kannst das Vorlesen jederzeit anhalten, beschleunigen und verlangsamen. Weitere Informationen hier.
Ist Scaling Big Data with Hadoop and Solr - Second Edition als Online-PDF/ePub verfügbar?
Ja, du hast Zugang zu Scaling Big Data with Hadoop and Solr - Second Edition von Hrishikesh Vijay Karambelkar im PDF- und/oder ePub-Format sowie zu anderen beliebten Büchern aus Computer Science & Programming in Java. Aus unserem Katalog stehen dir über 1 Million Bücher zur Verfügung.

Information

Jahr
2015
ISBN
9781783553402

Scaling Big Data with Hadoop and Solr Second Edition


Table of Contents

Scaling Big Data with Hadoop and Solr Second Edition
Credits
About the Author
About the Reviewers
www.PacktPub.com
Support files, eBooks, discount offers, and more
Why subscribe?
Free access for Packt account holders
Preface
What this book covers
What you need for this book
Who this book is for
Conventions
Reader feedback
Customer support
Downloading the example code
Errata
Piracy
Questions
1. Processing Big Data Using Hadoop and MapReduce
Apache Hadoop's ecosystem
Core components
Understanding Hadoop's ecosystem
Configuring Apache Hadoop
Prerequisites
Setting up ssh without passphrase
Configuring Hadoop
Running Hadoop
Setting up a Hadoop cluster
Common problems and their solutions
Summary
2. Understanding Apache Solr
Setting up Apache Solr
Prerequisites for setting up Apache Solr
Running Apache Solr on jetty
Running Solr on other J2EE containers
Hello World with Apache Solr!
Understanding Solr administration
Solr navigation
Common problems and solutions
The Apache Solr architecture
Configuring Solr
Understanding the Solr structure
Defining the Solr schema
Solr fields
Dynamic fields in Solr
Copying the fields
Dealing with field types
Additional metadata configuration
Other important elements of the Solr schema
Configuration files of Apache Solr
Working with solr.xml and Solr core
Instance configuration with solrconfig.xml
Understanding the Solr plugin
Other configuration
Loading data in Apache Solr
Extracting request handler – Solr Cell
Understanding data import handlers
Interacting with Solr through SolrJ
Working with rich documents (Apache Tika)
Querying for information in Solr
Summary
3. Enabling Distributed Search using Apache Solr
Understanding a distributed search
Distributed search patterns
Apache Solr and distributed search
Working with SolrCloud
Why ZooKeeper?
The SolrCloud architecture
Building an enterprise distributed search using SolrCloud
Setting up SolrCloud for development
Setting up SolrCloud for production
Adding a document to SolrCloud
Creating shards, collections, and replicas in SolrCloud
Common problems and resolutions
Sharding algorithm and fault tolerance
Document Routing and Sharding
Shard splitting
Load balancing and fault tolerance in SolrCloud
Apache Solr and Big Data – integration with MongoDB
What is NoSQL and how is it related to Big Data?
MongoDB at glance
Installing MongoDB
Creating Solr indexes from MongoDB
Summary
4. Big Data Search Using Hadoop and Its Ecosystem
Understanding NoSQL
Working with the Solr HDFS connector
Big data search using Katta
How Katta works?
Setting up the Katta cluster
Creating Katta indexes
Using Solr 1045 Patch – map-side indexing
Using Solr 1301 Patch – reduce-side indexing
Distributed search using Apache Blur
Setting up Apache Blur with Hadoop
Apache Solr and Cassandra
Working with Cassandra and Solr
Single node configuration
Integrating with multinode Cassandra
Scaling Solr through Storm
Getting along with Apache Storm
Advanced analytics with Solr
Integrating Solr and R
Summary
5. Scaling Search Performance
Understanding the limits
Optimizing search schema
Specifying default search field
Configuring search schema fields
Stop words
Stemming
Index optimization
Limiting indexing buffer size
When to commit changes?
Optimizing index merge
Optimize option for index merging
Optimizing the container
Optimizing concurrent clients
Optimizing Java virtual memory
Optimizing search runtime
Optimizing through search query
Filter queries
Optimizing the Solr cache
The filter cache
The query result cache
The document cache
The field value cache
The lazy field loading
Optimizing Hadoop
Monitoring Solr instance
Using SolrMeter
Summary
A. Use Cases for Big Data Search
E-Commerce websites
Log management for banking
The problem
How can it be tackled?
High-level design
Index

Scaling Big Data with Hadoop and Solr Second Edition

Copyright © 2015 Packt Publishing
All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews.
Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book.
Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information.
First published: August 2013
Second edition: April 2015
Production reference: 1230415
Published by Packt Publishing Ltd.
Livery Place
35 Livery Street
Birmingham B3 2PB, UK.
ISBN 978-1-78355-339-6
www.packtpub.com

Credits

Author
Hrishikesh Vijay Karambelkar
Reviewers
Ramzi Alqrainy
Walt Stoneburner
Ning Sun
Ruben Teijeiro
Commissioning Editor
Kartikey Pandey
Acquisition Editor
Nikhil Chinnari
Reshma Raman
Content Development Editor
Susmita Sabat
Technical Editor
Aman Preet Singh
Copy Editors
Sonia Cheema
Tani Kothari
Project Coordinator
Milton Dsouza
Proofreader
Simran Bhogal
Safis Editing
Indexer
Mariammal Chettiyar
Production Coordinator
Arvindkumar Gupta
Co...

Inhaltsverzeichnis