What needs improvement with Spark SQL?

Spark SQL is a Spark module for structured data processing. Unlike the basic Spark RDD API, the interfaces provided by Spark SQL provide Spark with more information about the structure of both the data and the computation being performed. There are several ways to interact with Spark SQL including SQL and the Dataset API. When computing a result the same execution engine is used, independent of which API/language you are using to express the computation. This unification means that developers...

Download Spark SQL Report Read more

Related Q&As

Aug 18, 2023

What is your experience regarding pricing and costs for Spark SQL?

Nov 23, 2023

What do you like most about Spark SQL?

SurjitChoudhury Data engineer at Cocos pt · Answer 1 · 2023-11-23T15:19:35Z

In terms of improvement, the only thing that could be enhanced is the stability aspect of Spark SQL. There could be additional features that I haven't explored but the current solution for working with databases seems effective. I haven't worked extensively with all components, so there might be untapped features that could enhance the solution's value.

Slaven Batnozic CTO at Dokument IT d.o.o. · Answer 2 · 2023-08-18T08:37:21Z

I'm using DBeaver to connect Spark with external tools. I've experienced some incompatibilities when using the Delta Lake format. It is compatible when you're using Databricks on the cloud, but when I'm using Spark on-premise, there are some incompatibility issues. We expect interactive queries with Dremio to provide better results. We issue a query but see that it's a batch process in the background. The documentation is also limited, especially in the setup for Thrift servers.

Aria Amini Data Engineer at Behsazan Mellat · Answer 3 · 2023-07-26T11:55:00Z

It would be useful if Spark SQL integrated with some data visualization tools. For example, we could integrate Spark SQL with Tableau for data visualization.

Sahil Taneja Principal Consultant/Manager at Tenzing · Answer 4 · 2023-05-05T08:54:14Z

Spark SQL can improve the documentation they have provided. It can be a bit unclear at times. They could improve the documentation a bit more so that we can understand it more easily. Moreover, they could improve SparkUI to have more advanced versions of the performance and the queries and all.

Lucas Dreyer Data Engineer at BBD · Answer 5 · 2023-01-04T13:37:06Z

It takes a bit of time to get used to using this solution versus Panda as it has a steep learning curve. You need quite a high level of skill with SQL in general to use this solution. If SQL is not someone's primary language, they might find it difficult to get used to. This solution could be improved if there was a bridge between Panda and Spark SQL such as translating from Panda operations to SQL and then working with those queries that are generated. In a future release, it would be useful to have a real time dashboard versus batch updates to Power BI.

score 0 · Answer 6 · 2022-11-22T13:27:47Z

It would be beneficial for aggregate functions to include a code block or toolbox that explains calculations or supported conditional statements. Multiple functions come within an aggregate so it is important to understand them. When you are trying to do something new, it would be easier and quite unique to get information within the solution rather than having to search the web. For example, once you select an aggregate it tells you what type of functions the solution can perform and includes a code block explaining its calculations. Or, a certain conditional statement gives you a second option or explains other types of statements the solution performs as part of a rule-level function.

Mahdi Sharifmousavi Lecturer at Amirkabir University of Technology · Answer 7 · 2022-08-10T11:49:13Z

There are many inconsistencies in syntax for the different querying tasks like selecting columns and joining between two tables so I'd like to see a more consistent syntax. Notations should be unified for all tasks within Spark SQL.

reviewer1724670 Engineering Manager/Solution architect at Provectus · Answer 8 · 2021-12-02T15:07:38Z

reviewer1724670

Engineering Manager/Solution architect at Provectus

Vendor

Dec 2, 2021

This solution could be improved by adding monitoring and integration for the EMR.

reviewer1488372 Associate Manager at a consultancy with 501-1,000 employees · Answer 9 · 2021-05-29T10:04:10Z

reviewer1488372

Associate Manager at a consultancy with 501-1,000 employees

Real User

May 29, 2021

There should be better integration with other solutions.

score 0 · Answer 10 · 2020-09-27T04:10:00Z

Being a new user, I am not able to find out how to partition it correctly. I probably need more information or knowledge. In other database solutions, you can easily optimize all partitions. I haven't found a quicker way to do that in Spark SQL. It would be good if you don't need a partition here, and the system automatically partitions in the best way. They can also provide more educational resources for new users.

Piotr Kalanski Cloud Team Leader at TCL · Answer 11 · 2020-04-26T06:32:00Z

I would like to have the ability to process data without the overhead. To use the same API to process both terabytes data and be able to process one GB of data.

score 0 · Answer 12 · 2020-03-18T06:06:00Z

Anything to improve the GUI would be helpful. We have experienced a lot of issues, but nothing in the production environment.

DulalMali Data Analytics Practice head at bse · Answer 13 · 2020-02-09T08:17:05Z

The service is complex. This is due to the fact that it's a combination of a lot of technology. The solution needs to include graphing capabilities. Including financial charts would help improve everything overall.

score 0 · Answer 14 · 2019-07-16T05:40:00Z

it_user986637

Project Manager - Senior Software Engineer at a tech services company with 11-50 employees

Real User

Jul 16, 2019

In the next release, maybe the visualization of some command-line features could be added.