Five Hard-Won Lessons Using Hive
A data engineer shares five practical lessons and performance tips for working with Apache Hive, focusing on common pitfalls and optimizations.
A data engineer shares five practical lessons and performance tips for working with Apache Hive, focusing on common pitfalls and optimizations.
Fixing MongoDB Connector for Hadoop authentication errors by granting the clusterManager role to the user.
An explanation of Microsoft Azure HDInsights, a managed Apache Hadoop service for processing big data on Azure.
Final tutorial on analyzing airline data with Hadoop using Hive for SQL queries and Pig for scripting, covering setup and basic analytics.
Explores how the demand for big data skills in industry is draining talent from academic science, threatening research.
A tutorial on using Apache Hive to create tables and views from data loaded into a Hadoop cluster, continuing a multi-part series.
Explains how to parallelize QR decomposition for linear models on big data using R's biglm package and incremental merging.
A practical guide introducing Hadoop's ecosystem and setting up a proof-of-concept cluster on Amazon EC2 using Cloudera for big data processing.
A guide to installing and using R on Amazon EC2 instances to overcome in-memory limitations for big data analysis.
Announcement for DevNexus 2013, a Java/JVM technology conference in Atlanta, featuring sessions on cloud, mobile, web, and more.