Apache Hadoop is a proven platform for long-term storage and archiving of structured and unstructured data. Related ecosystem tools, such as Apache Flume and Apache Sqoop, allow users to easily ingest structured and semi-structured data without requiring the creation of custom code. Unstructured data, however, is a more challenging subset of data that typically lends itself to batch-ingestion methods. Although such methods are suitable for many use cases,
Cloudera Enterprise 5.8 is now generally available (comprising CDH 5.8, Cloudera Manager 5.8, and Cloudera Navigator 2.7).
Cloudera is excited to announce the general availability of Cloudera Enterprise 5.8! Main highlights of this release include Impala read/write support on Amazon S3, a redesigned SQL query editor GUI, the expansion of role-based access control functionality to Cloudera Search, and the GA of Cloudera Navigator Optimizer to facilitate and optimize workload migrations.
This new release includes, among other things, support for “slicing and dicing” workloads by user/application/report, workload breakdown by similar queries, and alerts for Apache Hive and Apache Impala (incubating) best practices.
Cloudera Navigator Optimizer enables database architects and database administrators (DBAs) to gain in-depth understanding of their SQL workloads running in data warehouse environments or on Apache Hadoop. Navigator Optimizer makes planning offload projects more predictable by assessing risk and reducing development costs.
This new release contains, among other things, support for usage-based billing, deployments to Microsoft Azure, and deployments across providers or regions.
Cloudera Director is a manifestation of Cloudera’s commitment to provide a simple and reliable way to deploy, scale, and manage Apache Hadoop in the cloud of your choice. Cloudera Director enables you to deploy production-ready clusters for big data applications and successfully run workloads in the cloud. With Cloudera Director 2.1,
Learn how analyzing stats from professional sports leagues is an instructive use case for data analytics using Apache Spark with SQL. Covered in this installment: data exploration with Apache Impala (incubating) and Hue.
In Part 1 of this series, I introduced the topic of using fantasy sports analytics as an instructive use case for exploring the Apache Hadoop ecosystem. In that installment, we focused on data processing by taking a collection of data from Basketball-Reference.com and enriching it with z-scores and normalized z-scores to analyze the relative value of NBA players.