Category Archives: Data Science

How-to: Use MADlib Pre-built Analytic Functions with Impala

Categories: Data Science Guest Impala

Thanks to Victor Bittorf, a visiting graduate computer science student at Stanford University, for the guest post below about how to use the new prebuilt analytic functions for Cloudera Impala.

Cloudera Impala is an exciting project that unlocks interactive queries and SQL analytics on big data. Over the past few months I have been working with the Impala team to extend Impala’s analytic capabilities. Today I am happy to announce the availability of pre-built mathematical and statistical algorithms for the Impala community under a free open-source license.

Read more

Customer Spotlight: Persado Makes Marketing a Data Science

Categories: Data Science Hadoop Use Case

It’s common to hear people describe themselves as being “left-brained” or “right-brained” based on their tendency to be more logical and mathematically driven (left-brained), or, conversely, to be intuitive and creatively driven (right-brained). For example, people who prefer math over art are often considered left-brained. People who get a higher verbal score on their SATs than for math are often considered right-brained.

In general, language and creative writing are considered right-brained exercises. Many people also associate marketing and advertising as a right-brained function,

Read more

Meet the Project Founder: Josh Wills

Categories: Data Science Hadoop MapReduce Meet the Engineer

In this installment of “Meet the Project Founder,” we speak with Josh Wills (@josh_wills), Cloudera’s Senior Director of Data Science and founder of Apache Crunch and Cloudera ML.

What led you to your project idea(s)?
When I first started at Cloudera in 2011, I had a fairly vague job description, no real responsibilities, and wasn’t all that familiar with the Apache Hadoop stack, so I started working on various pet projects in order to learn more about the tools and the use cases in domains like healthcare and energy.

Read more

Get Hired as a Certified Data Scientist

Categories: Data Science Training

To paraphrase Nate Silver: “There is lots of data coming. Who will speak for all this data?”

Nearly every day, I read new articles about how Big Data is “changing everything.” Data scientists are unlocking new approaches that help researchers find the cure for cancer, banks fight fraud, the police fight drug-related crimes, and fantasy sports leaguers fight each other.

It seems like all I need is an analytics platform like Apache Hadoop and a big pile of data,

Read more

Myrrix Joins Cloudera to Bring "Big Learning" to Hadoop

Categories: Data Science Hadoop Mahout

What a short, strange trip it’s been. Just a year ago, I founded Myrrix in London’s Silicon Roundabout to commercialize large-scale machine learning based on Apache Hadoop and Apache Mahout. It’s been a busy scramble, building software and proudly watching early customers get real, big data-sized machine learning into production.

And now another beginning: Myrrix has a new home in Cloudera. I’m excited to join as Director of Data Science in London,

Read more