#SparkSQL
June 29, 2025 at 9:19 AM
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations www.infoq.com/articles/rep...
Building Reproducible ML Systems with Apache Iceberg and SparkSQL: Open Source Foundations
Traditional data lakes aregreat for storing massive amounts of stuff, but they're terrible at the transactional guarantees and versioning that ML workloads desperately need. Apache Iceberg and…
www.infoq.com
August 3, 2025 at 3:05 PM
🧠 Python & Spark SQL Transforms

Use Python for custom logic or Spark SQL for distributed processing. Mix both for maximum control.

Build fast. Iterate faster.

#PythonETL #SparkSQL #DataTransformation #DataPipelines #ETL #OpenSource #OpenETL #DataOmni #BigData #FlexibleEngineering #BuildInPublic
July 16, 2025 at 3:30 AM
We’ll be at the #Databricks Data + AI Summit in SF next week (6/9–12).

If you’re around and want to chat about how incremental computing can make your #SparkSQL workloads go from hours to seconds — let’s connect.

Grab some time here: calendly.com/matt-feldera...

#DataAISummit #DataEngineering
June 5, 2025 at 8:47 PM
Largely Ruby. I worked with some data engineering tools last year though - SparkSQL, Athena, PrestoDB, TinyBird and ClickHouse. Plus an internal data tool.
November 26, 2024 at 10:58 PM
I like it* for converting between t-sql and sparksql. Vibe-coding in Powershell generally sends me in a loop where I could have figured it out myself in the same time. I gave up on DAX a long time ago.

*it = the specific LLM we are allowed at work.
December 2, 2025 at 3:48 PM
🚀 Working with PySpark SQL? Here's a quick and powerful example!

You can query DataFrames using SQL syntax in Spark — great for teams coming from SQL backgrounds.

#PySpark #BigData #SparkSQL #DataEngineering #ETL #ApacheSpark #SQL #DataScience #XavierDataTech
June 28, 2025 at 8:57 PM
What languages can be used in Fabric Notebooks?
Microsoft Fabric Notebooks support:
🔹 PySpark
🔹 Spark (Scala)
🔹 SparkSQL
🔹 SparkR (R)
🔹 HTML
#MicrosoftFabric #FabricNotebooks #PySpark #SparkSQL #SparkR #Scala #BigData #DataEngineering #DataScience #OneLake #FabricCommunity #DataPlatform #DP700
July 28, 2025 at 3:07 AM
I’ve tried to use AI for code translation, TSQL to SparkSQL, and even when it does do it well it usually takes a lot of hoop-jumping because it rarely manages to consume more than 50 lines of code without getting lost.

I finally got off my ass and just learned SparkSQL.
February 11, 2026 at 6:24 PM
I think Eugene is spot on. In Fabric, MLVs are refreshed on a schedule (currently only full refresh is supported). I think of them the same as Data pipeline, but written in SparkSQL instead of using GUI
July 19, 2025 at 8:51 PM
5/5
Finally, a prototype implementation demonstrates the practical impact on query evaluation. Early experiments show potential for large performance gains on difficult queries in standard database systems (e.g., PostgreSQL, SparkSQL)

Check it out: arxiv.org/pdf/2412.11669
arxiv.org
December 20, 2024 at 8:40 AM
Just wait until they found that scalable, powerful native #SQL engines in #RDBMS that scale better and use CPU much more efficient, than #SparkSQL. 5-10 years down the line their code will run on #Oracle, #Snowflake or the likes.
June 28, 2025 at 4:06 PM
OpenMLDB can handle about 12,500 queries per second with sub‑millisecond latency, beating SparkSQL and ClickHouse by 23× and PostgreSQL/MySQL by 3.57×, according to recent benchmarks. Read more: https://getnews.me/openmldb-boosts-real-time-sql-ml-query-performance/ #openmldb #sql #ml
September 22, 2025 at 8:17 AM
Great talk by Binwei Yang on Apache Gluten last week.

youtu.be/GWTj3INSzPg?...

Apache Gluten moves execution of spark operators to native backend like Velox, accelerating query performance.
It has basic iceberg support too!
github.com/apache/incub...
Big Data Bellevue: Apache Gluten: Accelerating SparkSQL with Spark on Velox
YouTube video by BDB
youtu.be
January 19, 2025 at 2:06 AM
Azure Bigdata Specialist

Job title: Azure Bigdata Specialist Company: PradeepIT Job description: About the job Azure Bigdata Specialist Job Description Overall, 4 to 8 years of experience in IT Industry. Min 4..., Python, SparkSQL, Scala, Azure Blob Storage. Experience in Real-Time Data Processing…
Azure Bigdata Specialist
Job title: Azure Bigdata Specialist Company: PradeepIT Job description: About the job Azure Bigdata Specialist Job Description Overall, 4 to 8 years of experience in IT Industry. Min 4..., Python, SparkSQL, Scala, Azure Blob Storage. Experience in Real-Time Data Processing using Apache Kafka/EventHub/IoT... Expected salary: Location: Bangalore, Karnataka Job date: Wed, 26 Feb 2025 08:02:25 GMT Apply for the job now!
findsuperdeals.shop
May 16, 2025 at 11:45 AM
Spark RDDs Vs DataFrames vs SparkSQL – Part 3 : Web Server Log Analysis

#python
Spark RDDs Vs DataFrames vs SparkSQL – Part 3 : Web Server Log Analysis
This is the third tutorial on the Spark RDDs Vs DataFrames vs SparkSQL blog post series. The first one is available here. In the first part, we saw how to…
datascienceplus.com
August 2, 2026 at 6:00 PM
Spark RDDs Vs DataFrames vs SparkSQL – Part 2 : Working With Multiple Tables

#python
Spark RDDs Vs DataFrames vs SparkSQL – Part 2 : Working With Multiple Tables
This is the second tutorial on the Spark RDDs Vs DataFrames vs SparkSQL blog post series. The first one is available at DataScience+. In the first part, I showed how…
datascienceplus.com
July 27, 2026 at 6:00 PM
There's just something beautiful about watching Spark pods spin up in k8s. 😍
December 6, 2024 at 6:54 PM
At @hadoopsummit? @ted_dunning #MapR speaking #SparkSQL vs @apachedrill in 30 min at Ballroom C! #hs16sj
November 23, 2024 at 5:32 PM
continue to love an appreciate your site so much. curious how you decided which new engine to add? Maybe my experience is biased since I'm a data engineer in a corporate environment, but I was happy to see BigQuery and was hoping the next add would be Snowflake or SparkSQL
January 15, 2026 at 12:58 AM
Convert SQL to Pyspark or SparkSQL
stackoverflow.com
November 25, 2024 at 6:40 PM
Today in Big Data: Use your protobuf definitions to define a SparkSQL schema.

https://gist.github.com/ebenoist/80542b95d1c5d0c0aaabaa37ee154922
November 12, 2024 at 3:35 AM
plugged into any standard SQL engine. Our system prototype currently supports four different SQL engines (DuckDB, PostgreSQL, SparkSQL, and AnalyticDB from Alibaba Cloud), and our experiments show that Yannakakis+ is able to deliver better performance [5/6 of https://arxiv.org/abs/2504.03279v1]
April 7, 2025 at 5:55 AM