#ApacheSpark
Need a distraction? Come and join me for hacking on @ApacheSpark today: www.youtube.com/watch?v=h1fD... / twitch.tv/holdenkarau
Spark PR Follow Up (splitting) and SPIP progress
YouTube video by Holden Karau
www.youtube.com
January 12, 2026 at 8:09 PM
Come and join me for a quick pre-holiday @ApacheSpark stream: www.twitch.tv/holdenkarau hacking on filter pushdown regression.
holdenkarau - Twitch
Pre-Holiday Spark PR Follow Up Stream: Updating Pushdown Logic
www.twitch.tv
December 18, 2025 at 9:51 PM
Come say hi if you see me at #current23 love to chat with folks about @ApacheSpark , @ApacheIceberg , @dask_dev or @raydistributed 😊 #leathertuesday
September 26, 2023 at 7:00 PM
Building #ApacheSpark 4.0.0.dev0 (master) takes 1h on mac mini / Apple M2 Pro 😱

Would M4 make it faster? Half the time, perhaps?
January 5, 2025 at 2:59 PM
One of my resolutions for 2025: champion #uv as the only #Python project manager of choice. It is so pleasant to use!

#ApacheSpark Connect requires some extra deps. No need for a venv, just "uv run --with".

Make uv your Python project manager in 2025 🙏

➡️ docs.astral.sh/uv/

HNY 🥂🥳
January 1, 2025 at 2:37 PM
June 14, 2026 at 1:25 PM
#CaseStudy - #Lyft rearchitected its ML platform, LyftLearn, into a hybrid system!

Offline workloads now run on AWS SageMaker, while Kubernetes continues to power online model serving.

The result❓ Read #InfoQ and find out 👉 bit.ly/4s4mf9j

#SoftwareArchitecture #AI #ML #ApacheSpark #Kubernetes
December 19, 2025 at 7:01 AM
Come and join me for more hacking on @apachespark today -- www.youtube.com/watch?v=u5F_... I'm going to be working on github.com/apache/spark...
Spark PR adventures: Determinism Issues (maybe?)
YouTube video by Holden Karau
www.youtube.com
November 13, 2025 at 6:59 PM
There are 6+ different SparkSessions in the upcoming #ApacheSpark 4.0.0 (snapshot) due to #SparkConnect 🤯

1️⃣ api.SparkSession
2️⃣ SparkSession
3️⃣ Spark Connect-aware SparkSession

Hope to get it all right in "The Internals of Spark Connect"...soon 🫣

➡️ books.japila.pl/spark-connec...
December 22, 2024 at 12:49 PM
December 10, 2025 at 3:46 PM
96% fewer out-of-memory (OOM) failures!

#Pinterest boosted #ApacheSpark reliability by:
✅ Enhanced observability
✅ Config tuning
✅ Automatic memory retries
The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

⇨ bit.ly/4e9n1gD

#InfoQ
April 8, 2026 at 8:49 AM
🚀 Working with PySpark SQL? Here's a quick and powerful example!

You can query DataFrames using SQL syntax in Spark — great for teams coming from SQL backgrounds.

#PySpark #BigData #SparkSQL #DataEngineering #ETL #ApacheSpark #SQL #DataScience #XavierDataTech
June 28, 2025 at 8:57 PM
Correct. That's to learn the internals of #ApacheSpark using JDWP.
December 29, 2024 at 5:56 PM
🔥 Big data, big possibilities! 🔥 Apache Spark is at the heart of large-scale data analytics, and Microsoft Fabric makes it easier than ever. 🎓 Get started today: https://sbee.link/x7djeuf6qv

#MicrosoftLearn #ApacheSpark #BigData #MicrosoftFabric #TheTrustedAdvisor
May 15, 2025 at 8:09 AM
Come and join me for proposing a large(ish) @ApacheSpark change and see how we run the SPIP process -- www.youtube.com/watch?v=bC2y... :D
Very mini stream: Starting discussion for an SPIP around transpilation
YouTube video by Holden Karau
www.youtube.com
December 19, 2025 at 7:29 PM
As #MLflow maintainers, we were excited to share MLflow's #GenAI features with Bilibili’s team! 🚀

With their strong #opensource culture + #ApacheSpark use, MLflow is a natural fit to support their growing ML and GenAI initiatives. 🔥
April 29, 2025 at 1:28 PM
We argue that existing data transfer tools fall short in today's diverse environments. We introduce XDBC, a framework for fast and scalable data transfer across heterogeneous systems and infrastructures, whether you're working with #PostgreSQL, #Pandas, #ApacheSpark, or anything in between.
June 20, 2025 at 5:51 PM
Discover how Decathlon, one of the world’s leading sports retailers, adopted the #opensource library Polars to optimize its data workflows.

By migrating from #ApacheSpark to #Polars for small input datasets, Decathlon achieved:
• Significant speed
• Cost savings

👉 bit.ly/4atNCTY

#InfoQ #AI
December 22, 2025 at 1:34 PM