#SparkSQL
Spark RDDs Vs DataFrames vs SparkSQL – Part 5: Using Functions

#python
Spark RDDs Vs DataFrames vs SparkSQL – Part 5: Using Functions
This is the fifth tutorial on the Spark RDDs Vs DataFrames vs SparkSQL blog post series. The first one is available here. In the first part, we saw how to…
datascienceplus.com
September 25, 2026 at 10:00 PM
Spark RDDs Vs DataFrames vs SparkSQL – Part 3 : Web Server Log Analysis

#python
Spark RDDs Vs DataFrames vs SparkSQL – Part 3 : Web Server Log Analysis
This is the third tutorial on the Spark RDDs Vs DataFrames vs SparkSQL blog post series. The first one is available here. In the first part, we saw how to…
datascienceplus.com
August 2, 2026 at 6:00 PM
Spark RDDs Vs DataFrames vs SparkSQL – Part 2 : Working With Multiple Tables

#python
Spark RDDs Vs DataFrames vs SparkSQL – Part 2 : Working With Multiple Tables
This is the second tutorial on the Spark RDDs Vs DataFrames vs SparkSQL blog post series. The first one is available at DataScience+. In the first part, I showed how…
datascienceplus.com
July 27, 2026 at 6:00 PM
I’ve tried to use AI for code translation, TSQL to SparkSQL, and even when it does do it well it usually takes a lot of hoop-jumping because it rarely manages to consume more than 50 lines of code without getting lost.

I finally got off my ass and just learned SparkSQL.
February 11, 2026 at 6:24 PM
continue to love an appreciate your site so much. curious how you decided which new engine to add? Maybe my experience is biased since I'm a data engineer in a corporate environment, but I was happy to see BigQuery and was hoping the next add would be Snowflake or SparkSQL
January 15, 2026 at 12:58 AM
🚀 New Lab Replay: Using Delta Tables in Apache Spark (Microsoft Fabric)
🎥 Watch the full session:
👉 www.youtube.com/live/gT21FS8...

#MicrosoftFabric #DeltaTables #ApacheSpark #DeltaLake #DP600 #DP700 #Lakehouse #DataEngineering #BigData #ACID #TimeTravel #SparkSQL #PySpark #MicrosoftLearn
December 8, 2025 at 12:30 PM
I like it* for converting between t-sql and sparksql. Vibe-coding in Powershell generally sends me in a loop where I could have figured it out myself in the same time. I gave up on DAX a long time ago.

*it = the specific LLM we are allowed at work.
December 2, 2025 at 3:48 PM
OpenMLDB can handle about 12,500 queries per second with sub‑millisecond latency, beating SparkSQL and ClickHouse by 23× and PostgreSQL/MySQL by 3.57×, according to recent benchmarks. Read more: https://getnews.me/openmldb-boosts-real-time-sql-ml-query-performance/ #openmldb #sql #ml
September 22, 2025 at 8:17 AM
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations
Traditional data lakes aregreat for storing massive amounts of stuff, but they're terrible at the transactional guarantees and versioning that ML workloads desperately need. Apache Iceberg and…
www.infoq.com
August 6, 2025 at 2:28 AM
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations www.infoq.com/articles/rep...
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations
Traditional data lakes aregreat for storing massive amounts of stuff, but they're terrible at the transactional guarantees and versioning that ML workloads desperately need. Apache Iceberg and…
www.infoq.com
August 3, 2025 at 3:05 PM
What languages can be used in Fabric Notebooks?
Microsoft Fabric Notebooks support:
🔹 PySpark
🔹 Spark (Scala)
🔹 SparkSQL
🔹 SparkR (R)
🔹 HTML
#MicrosoftFabric #FabricNotebooks #PySpark #SparkSQL #SparkR #Scala #BigData #DataEngineering #DataScience #OneLake #FabricCommunity #DataPlatform #DP700
July 28, 2025 at 3:07 AM
I think Eugene is spot on. In Fabric, MLVs are refreshed on a schedule (currently only full refresh is supported). I think of them the same as Data pipeline, but written in SparkSQL instead of using GUI
July 19, 2025 at 8:51 PM
🧠 Python & Spark SQL Transforms

Use Python for custom logic or Spark SQL for distributed processing. Mix both for maximum control.

Build fast. Iterate faster.

#PythonETL #SparkSQL #DataTransformation #DataPipelines #ETL #OpenSource #OpenETL #DataOmni #BigData #FlexibleEngineering #BuildInPublic
July 16, 2025 at 3:30 AM
June 29, 2025 at 9:19 AM
🚀 Working with PySpark SQL? Here's a quick and powerful example!

You can query DataFrames using SQL syntax in Spark — great for teams coming from SQL backgrounds.

#PySpark #BigData #SparkSQL #DataEngineering #ETL #ApacheSpark #SQL #DataScience #XavierDataTech
June 28, 2025 at 8:57 PM
Just wait until they found that scalable, powerful native #SQL engines in #RDBMS that scale better and use CPU much more efficient, than #SparkSQL. 5-10 years down the line their code will run on #Oracle, #Snowflake or the likes.
June 28, 2025 at 4:06 PM
We’ll be at the #Databricks Data + AI Summit in SF next week (6/9–12).

If you’re around and want to chat about how incremental computing can make your #SparkSQL workloads go from hours to seconds — let’s connect.

Grab some time here: calendly.com/matt-feldera...

#DataAISummit #DataEngineering
June 5, 2025 at 8:47 PM
Azure Bigdata Specialist

Job title: Azure Bigdata Specialist Company: PradeepIT Job description: About the job Azure Bigdata Specialist Job Description Overall, 4 to 8 years of experience in IT Industry. Min 4..., Python, SparkSQL, Scala, Azure Blob Storage. Experience in Real-Time Data Processing…
Azure Bigdata Specialist
Job title: Azure Bigdata Specialist Company: PradeepIT Job description: About the job Azure Bigdata Specialist Job Description Overall, 4 to 8 years of experience in IT Industry. Min 4..., Python, SparkSQL, Scala, Azure Blob Storage. Experience in Real-Time Data Processing using Apache Kafka/EventHub/IoT... Expected salary: Location: Bangalore, Karnataka Job date: Wed, 26 Feb 2025 08:02:25 GMT Apply for the job now!
findsuperdeals.shop
May 16, 2025 at 11:45 AM
#SparkSQL has been sent to try my patience. Just bloody work and stop complaining!
May 12, 2025 at 3:41 PM
"AI engineers should be able to analyze large sets of data and extract meaningful insights from them. This involves using big data tools such as SparkSQL, Apache Flink, Apache Arrow, and Google Cloud Platform to query and manipulate large datasets."
May 12, 2025 at 11:23 AM
plugged into any standard SQL engine. Our system prototype currently supports four different SQL engines (DuckDB, PostgreSQL, SparkSQL, and AnalyticDB from Alibaba Cloud), and our experiments show that Yannakakis+ is able to deliver better performance [5/6 of https://arxiv.org/abs/2504.03279v1]
April 7, 2025 at 5:55 AM