Try our new research platform with insights from 80,000+ expert users

Amazon EMR vs Cloudera Data Platform comparison

 

Comparison Buyer's Guide

Executive SummaryUpdated on Apr 1, 2025

Review summaries and opinions

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
 

ROI

Sentiment score
6.3
Companies using Amazon EMR often experience significant ROI, with savings up to 20% and substantial returns over on-premise systems.
Sentiment score
5.6
ROI experiences vary for entities using Cloudera, with some achieving 300% returns, influenced by Hadoop infrastructure and expectations.
 

Customer Service

Sentiment score
7.6
Amazon EMR support is generally proactive and efficient, but experiences vary, especially during open-source product integration.
Sentiment score
6.6
Cloudera's customer support varies widely, with mixed effectiveness ratings, reliant on community or paid services for complex issues.
They help with billing, cost determination, IAM properties, security compliance, and deployment and migration activities.
 

Scalability Issues

Sentiment score
7.8
Amazon EMR effectively scales to enterprise needs, with auto-scaling and adaptability, despite occasional peak demand resource allocation delays.
Sentiment score
6.4
Cloudera Data Platform is scalable but can face upgrade challenges; praised for handling large data, additional hardware may be needed.
Scalability can be provisioned using the auto-scaling feature, EC2 instances, on-demand instances, and storage locations like block storage, S3, or file storage.
For scalability, I rate Cloudera Data Platform at an eight out of ten as it is an on-premise solution.
 

Stability Issues

Sentiment score
8.1
Amazon EMR is generally stable and reliable, despite occasional data-related stability issues, with robust failover and monitoring features.
Sentiment score
7.1
Cloudera Data Platform is stable when configured correctly, with occasional HBase issues, achieving high user ratings despite setup complexities.
Regular updates, patch installations, monitoring, logging, alerting, and disaster recovery activities are crucial for maintaining stability.
 

Room For Improvement

Amazon EMR struggles with a steep learning curve, complex configurations, unpredictable costs, and needs enhancements in stability and support.
Cloudera Data Platform needs UI, security, and usability improvements, with challenges in integration, stability, and community resources.
There is room for improvement with respect to retries, handling the volume of data on S3 buckets, cluster provisioning, scaling, termination, security, and integration between services like S3, Glue, Lake Formation, and DynamoDB.
We aim to address these issues with a Kubernetes-based platform that will simplify the task of upgrading services.
 

Setup Cost

Amazon EMR's costs vary by resources used, with potential high monthly expenses, requiring careful management to prevent surprises.
Cloudera Data Platform offers per-node pricing, with free and paid tiers, and is cost-effective compared to Oracle.
Cost optimization can be achieved through instance usage, cluster sharing, and auto-scaling.
Initially, CDH had a straightforward pricing model based on nodes, but CDP includes factors like processors, cores, terabytes, and drives, making it difficult to calculate costs.
 

Valuable Features

Amazon EMR is scalable, easy to use, cost-effective, integrates well with Hadoop, and supports diverse analytics applications.
Cloudera Data Platform is favored for its scalability, open-source flexibility, easy deployment, and robust governance with seamless tool integration.
Amazon EMR helps in scalability, real-time and batch processing of data, handling efficient data sources, and managing data lakes, data stores, and data marts on file systems and in S3 buckets.
By using the Hadoop File System for distributed storage, we have 1.5 petabytes of physical storage with 500 terabytes of effective storage due to a replication factor of three.
 

Categories and Ranking

Amazon EMR
Average Rating
7.8
Reviews Sentiment
7.2
Number of Reviews
23
Ranking in other categories
Hadoop (3rd), Cloud Data Warehouse (12th)
Cloudera Data Platform
Average Rating
8.0
Reviews Sentiment
6.4
Number of Reviews
26
Ranking in other categories
Cloud Master Data Management (MDM) Solutions (10th), Data Management Platforms (DMP) (7th)
 

Featured Reviews

Prashant  Singh - PeerSpot reviewer
Seamless data integration enhances reporting efficiency and an easy setup
Amazon EMR has multiple connectors that can connect to various data sources. The service charges are based on processing only, depending on the resources used, which can help save money. It is easy to integrate with other services for storage, allowing data to be shifted to cheaper storage based on usage.
Miodrag-Stanic - PeerSpot reviewer
Distributed computing improves data processing while upgrade complexity needs addressing
There are challenges with upgrading or updating various services like Spark, Impala, and Hive on on-premise and bare metal solutions. We aim to address these issues with a Kubernetes-based platform that will simplify the task of upgrading services. We also wish to implement lakehouse capabilities with Iceberg or Delta Lake frameworks.
report
Use our free recommendation engine to learn which Hadoop solutions are best for your needs.
848,716 professionals have used our research since 2012.
 

Top Industries

By visitors reading reviews
Financial Services Firm
26%
Computer Software Company
14%
Educational Organization
8%
Manufacturing Company
7%
No data available
 

Company Size

By reviewers
Large Enterprise
Midsize Enterprise
Small Business
 

Questions from the Community

What do you like most about Amazon EMR?
Amazon EMR is a good solution that can be used to manage big data.
What is your experience regarding pricing and costs for Amazon EMR?
The cost of Amazon EMR is a little bit expensive, especially considering the support package, which includes a gold package.
What needs improvement with Amazon EMR?
Spark jobs take longer on Amazon EMR compared to previous experiences. This aspect could be improved to make them more efficient.
What do you like most about Hortonworks Data Platform?
Distributed computing, secure containerization, and governance capabilities are the most valuable features.
What is your experience regarding pricing and costs for Hortonworks Data Platform?
I haven't done a price analysis specifically for HDP. However, when it was first introduced as Hadoop 2.0, there were a few use cases where the price was quite high. It was particularly expensive f...
What needs improvement with Hortonworks Data Platform?
Since Cloudera acquired HDP, it's been bundled with CBH and HDP. However, the biggest challenge is cloud storage integration with Azure, GCP, and AWS. These platforms offer competitive storage solu...
 

Also Known As

Amazon Elastic MapReduce
No data available
 

Overview

 

Sample Customers

Yelp
Information Not Available
Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop. Updated: March 2025.
848,716 professionals have used our research since 2012.