#apache-spark
The rest of Palantir's stack is basically Apache Spark with a proprietary "Foundry Pipelines" sticker on it that forces you to write to their proprietary, locked-down git wrappers to write python 3.8 code.

Palantir Phonograph and Workshop are just GUI builders that feel like Microsoft Access.
June 19, 2026 at 8:03 PM
Sail

A drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.

github.com/lakehq/sail
April 29, 2026 at 12:45 AM
Check out the latest release of the Comet accelerator for Apache Spark

datafusion.apache.org/blog/2025/09...
Apache DataFusion Comet 0.10.0 Release - Apache DataFusion Blog
datafusion.apache.org
September 18, 2025 at 3:06 AM
DataFusion Comet 0.7.0 is now available in Maven. We'll be publishing a blog post next week with all the details.

The repo has been updated with the latest benchmark results. For single executor TPC-H @ 100 GB, we now see a 2.2x increase over Spark (up from 2x in 0.6.0).

github.com/apache/dataf...
GitHub - apache/datafusion-comet: Apache DataFusion Comet Spark Accelerator
Apache DataFusion Comet Spark Accelerator. Contribute to apache/datafusion-comet development by creating an account on GitHub.
github.com
March 19, 2025 at 5:11 PM
We have a position open in the Spark team at Apple, in our Cupertino, CA office. The role would include working on Apache DataFusion Comet.

jobs.apple.com/en-us/detail...
Senior Software Development Engineer (Apache Spark) - Apple Data Platform - Jobs - Careers at Apple
Apply for a Senior Software Development Engineer (Apache Spark) - Apple Data Platform job at Apple. Read about the role and find out if it’s right for you.
jobs.apple.com
April 2, 2025 at 5:28 PM
PySpark is a Python API for Apache Spark - and it lets you write Spark apps in Python. And PySpark is fast, as it distributes tasks across multiple machines. In this guide, Manish walks you through how to process data with Apache Spark & Python.
https://buff.ly/3KZe7lY
January 11, 2025 at 5:01 PM
he earliest commit we could find that used an coding agent was in Apache Spark back in November 2023: github.com/apache/spark...
[SPARK-46042][FOLLOWUP][CONNECT] Test and adapt to streaming RPC beha… · apache/spark@9c291e1
…vior change from grpc 1.56 to 1.59 ### What changes were proposed in this pull request? This is a followup to https://github.com/apache/spark/pull/43955 In grpc 1.56, when calling a server stre...
github.com
August 13, 2026 at 1:08 AM
"Introducing SedonaDB: A single-node analytical database engine with geospatial as a first-class citizen"

Built in Rust with Apache DataFusion

sedona.apache.org/latest/blog/...
Introducing SedonaDB: A single-node analytical database engine with geospatial as a first-class citizen - Apache Sedona
Apache Sedona is a cluster computing system for processing large-scale spatial data. Sedona extends existing cluster computing systems, such as Apache Spark, Apache Flink, and Snowflake, with a set of...
sedona.apache.org
September 24, 2025 at 9:20 PM
Want to read #ducklake from APACHE SPARK? Check this out: gist.github.com/hannes/395ac... #butdoesitscale #yesitdoes
pyspark-ducklake-2.py
GitHub Gist: instantly share code, notes, and snippets.
gist.github.com
July 1, 2025 at 11:23 AM
Algorithm
Show me post about

System design
Backend
Golang
Rust
Data Platform
Apache
Hadoop
Spark
Hudi
Deltalake
ScyllaDB
AWS
Cassandra
Kafka
Kubernetes
Docker
Jenkins
ArgoCI
November 9, 2024 at 5:38 PM
I’m excited to announce that I’ve joined Snowflake working OSS Apache Spark and I’ve written a post about what we’re doing careers.snowflake.com/us/en/blogar... :) (don’t worry I’m still working on Fight Health Insurance too thanks to coffee powered scaling)
Building Apache Spark™ in the Open at Snowflake
Why Holden Karau joined and what she's building at Snowflake.
careers.snowflake.com
February 12, 2026 at 2:55 PM
Gee, I wonder how they associated "spark" with machine learning. en.wikipedia.org/wiki/Apache_...
Apache Spark - Wikipedia
en.wikipedia.org
September 12, 2026 at 2:27 AM
parallel eyes
parallelize
parallel lies

(this bleet sponsored by Apache Spark)
April 20, 2023 at 6:07 PM
Apache Iceberg 1.11.0 is live!! 🧊 Unlocking Google KMS backed built-in encryption, faster GCP Lakehouse analytics queries. Read: goo.gle/apache-icebe...
goo.gle
Apache Iceberg 1.11.0 is here! Discover Spark 4.1 support, server-side scan planning, Google KMS table encryption, and faster storage analytics.
goo.gle
May 27, 2026 at 6:31 PM
The Spark Cassandra Connector has been donated to the project 💥

It can now be found at
github.com/apache/cassa... 👀
March 25, 2025 at 1:27 PM
If you're interested in learning more about accelerating Apache Spark with Apache DataFusion's Comet subproject, check out this talk I recently gave as part of CMU's Database Building Blocks Seminar Series.

We'd love for more people to try out Comet and give us feedback!

youtu.be/o59s0d3HE1k?...
Accelerating Apache Spark Workloads with Apache DataFusion Comet (Andy Grove)
YouTube video by CMU Database Group
youtu.be
October 25, 2024 at 4:15 PM
Want to get involved in open source (specifically Apache Spark)? We're running a Community Sprint in Seattle luma.com/rrfvx0ey for folks wanting to get started on Friday, March 13
Apache Spark™ Community Sprint · Luma
Apache Spark™ Community Sprint! Join us on March 17th (Tuesday) from 12:00-7:00 PM at the Snowflake Bellevue Office for a Spark community sprint! We'll spend…
luma.com
February 23, 2026 at 7:37 PM
Sparklyr 1.9.0 is out.

- Adds new Java folder for Spark 4.0.0 with updated code
- Adds new JAR file to handle Spark 4+
- Updates to different spots in the R code to start handling version 4
- Removes JARs using Scala 2.11

github.com/sparklyr/spa...

#rstats #spark #distributedComputing
GitHub - sparklyr/sparklyr: R interface for Apache Spark
R interface for Apache Spark. Contribute to sparklyr/sparklyr development by creating an account on GitHub.
github.com
March 24, 2025 at 2:59 PM
truly a rockstar building Apache Spark in the open. congrats all around!
I’m excited to announce that I’ve joined Snowflake working OSS Apache Spark and I’ve written a post about what we’re doing careers.snowflake.com/us/en/blogar... :) (don’t worry I’m still working on Fight Health Insurance too thanks to coffee powered scaling)
Building Apache Spark™ in the Open at Snowflake
Why Holden Karau joined and what she's building at Snowflake.
careers.snowflake.com
February 12, 2026 at 3:24 PM
On behalf of the DataFusion PMC, I'm excited to announce the release of version 0.11.0 of the Comet accelerator for Apache Spark!

datafusion.apache.org/blog/2025/10...
Apache DataFusion Comet 0.11.0 Release - Apache DataFusion Blog
datafusion.apache.org
October 22, 2025 at 2:21 PM
Yahoo optimizes data infrastructure on Google Cloud by using flexible VMs in Managed Service for Apache Spark clusters. 🌐⚙️ #Yahoo #GoogleCloud #BigData
How Yahoo Optimizes Apache Spark with Flexible VMs | Google Cloud Blog
Learn how Yahoo uses flexible VMs in Managed Service for Apache Spark to automatically handle capacity limits and reduce provisioning failures by 85%.
cloud.google.com
September 27, 2026 at 8:39 PM
Come and join me for a relaxing Friday afternoon of hacking on Apache Spark — twitch.tv/holdenkarau 👍
holdenkarau - Twitch
Holden is a transgender Canadian open source developer with a focus on Apache Spark and related
twitch.tv
November 7, 2025 at 9:35 PM
Amazon EMR introduces Long Term Support with Apache Spark 4.1

Amazon EMR introduces Long Term Support (LTS) releases, starting with emr-spark-8.1.0 and Apache Spark 4.1. With LTS, designated versions of the AWS runtime for Apache Spark receive 36 months of support. Amazon EMR pr...

#AWS #AmazonEmr
Amazon EMR introduces Long Term Support with Apache Spark 4.1
Amazon EMR introduces Long Term Support (LTS) releases, starting with emr-spark-8.1.0 and Apache Spark 4.1. With LTS, designated versions of the AWS runtime for Apache Spark receive 36 months of support. Amazon EMR provides LTS releases with fixes for critical and high severity security, bug, and data-corruption issues, subject to availability. LTS helps you run production Spark workloads on one release longer and upgrade on your own schedule, at no additional cost. This release adds full support for Apache Iceberg v3, bringing new geospatial, high-precision timestamp, and schema-evolution capabilities to your tables. Spark SQL queries can reference catalogs by name, including cross-account and Amazon S3 Tables catalogs, and automatically detect Apache Iceberg, Delta Lake, and Apache Hudi table formats, without registering each catalog in your Spark configuration. Fine-grained access control now covers more Apache Iceberg operations and the Delta Lake VACUUM operation, so you can apply column-level and row-level permissions to a wider set of jobs. Amazon EMR on EKS clusters now support Spark Connect endpoints with token-based authentication. emr-spark-8.1.0 is available in all AWS Regions where Amazon EMR is available, across Amazon EMR on EC2, Amazon EMR on EKS, and Amazon EMR Serverless. To learn more, see the https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-spark810-release.html and the https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-standard-support.html policy. To get started, create an EMR cluster or application with emr-spark-8.1.0 from the AWS Management Console, or use the https://docs.aws.amazon.com/emr/latest/ReleaseGuide/spark-upgrades.html to move existing applications to the release.
aws.amazon.com
September 23, 2026 at 12:05 AM
A gentle introduction to reasoning over large knowledge graph bases, extensive and complex ontologies, containing millions or billions of facts using Apache Cassandra + Apache Spark.

NORA: Scalable OWL reasoner …
onlinelibrary.wiley.com/doi/10.1002/...
November 23, 2024 at 11:19 AM