#PySpark
I wrote up a tutorial on an Introduction to PySpark
Intro to Data Analysis using PySpark
In this tutorial we will be exploring the functionality of PySpark on a World Population data set....
dev.to
January 21, 2025 at 8:14 PM
December 16, 2024 at 6:30 PM
48M14S by PySpark
6M07S by Python with #DuckDB

... that is a difference.
August 3, 2025 at 2:53 PM
Want to read #ducklake from APACHE SPARK? Check this out: gist.github.com/hannes/395ac... #butdoesitscale #yesitdoes
pyspark-ducklake-2.py
GitHub Gist: instantly share code, notes, and snippets.
gist.github.com
July 1, 2025 at 11:23 AM
A data scientist switching between pandas and pyspark multiple times for a single data pipeline and job is unhinged criminal behavior
August 19, 2026 at 1:25 PM
A DuckLake reader for PySpark in ~30 lines of code. Despite looking simple, still partitions work across Spark nodes. Zero external dependencies, just the DuckDB JDBC driver. Try that with your favorite lakehouse technology :-)
July 2, 2025 at 7:20 PM
PySpark is a Python API for Apache Spark - and it lets you write Spark apps in Python. And PySpark is fast, as it distributes tasks across multiple machines. In this guide, Manish walks you through how to process data with Apache Spark & Python.
https://buff.ly/3KZe7lY
January 11, 2025 at 5:01 PM
I'm super appreciative of @databard.bsky.social explaining PySpark in a way that doesn't assume any experience. So many things I wish I knew 4 months ago.
www.youtube.com/watch?v=2p2S...
ETL Approaches When Using Fabric PySpark - Jared Kuehn
YouTube video by Level Up Your Data
www.youtube.com
September 16, 2025 at 8:23 PM
Distributed Polars is 3x faster than Spark on the PDSH benchmark and up to 7.8x faster on individual queries.

Read the full benchmark post here: https:/pola.rs/posts/polars-pyspark-benchmarks/
June 16, 2026 at 2:00 PM
After many tests of Notebooks and PySpark Notebooks in #MicrosoftFabric, I am a bit losing scenarios for PySpark 😂. I still have some, but a lot can be handled directly with Python. Even bulk writing to Warehouse by SQL endpoint is slower than pushing Parquets and optimizing.
January 20, 2025 at 12:48 PM
For those who have issues with the PySpark schema definition... this might help: preetranjan.github.io/pyspark-sche...
Pyspark Schema Generator
preetranjan.github.io
March 11, 2025 at 9:23 AM
Pessoal, dou aula numa pós de engenharia de dados. Deixo as aulas disponíveis no Github pois minha meta é levar a educação a todas as camadas da sociedade.

github.com/laysabelici

Repositórios: Datas As a Service, Dw Redshift, Apache Beam, Pyspark, Databricks.

#bolhadev #dados
September 4, 2024 at 2:09 PM
dev.to/kingjotaro/s...

Configuração de um cluster pyspark em docker, pyspark é uma ferramenta de processamento de dados da família Apache Spark que da para brincar com multithread e processar grande quantidade de dados.

cc: @sseraphini.bsky.social, @samsantosb.bsky.social, @bolhadev.com #bolhadev
Setting Up a PySpark Cluster with Docker: Guide
Diving back into Python with a focus on expanding my knowledge in data processing, which has sparked...
dev.to
September 2, 2024 at 4:06 PM
O'Reilly Media - PySpark
"PySpark is the Python API for Apache Spark, a distributed framework for large-scale data analysis. It allows users to leverage the power of Spark using the Python programming language. PySpark"...

www.oreilly.com/search/skill...

===
#librecanada #linux #opensource
July 11, 2026 at 9:47 PM
Whooo hoooo

sqlserver —> glue/pyspark —> s3/parquet ran successfully!

Lots of fiddly things to look out for. But up and running!
March 5, 2026 at 4:28 PM
Olha a vaga de Engenharia de dados JR!
No Picpay e home office.
Requisitos: Graduação, SQL, Python, Pyspark, Cloud e DataOps. Bem de boas galera, nada fora do normal.
Por isso recomendo sempre, pegue SQL e refaça em Pyspark.

#vagasTI #BolhaDev #dados

picpay.com/oportunidade...
Central de Vagas PicPay: conheça e candidate-se
Saiba tudo sobre a vaga de Engenheiro de Dados Jr e candidate-se agora!
picpay.com
September 3, 2024 at 2:26 PM
Datalake, Databricks, Pyspark e Synapse
A galera tem mania de dar nomes de ninhas de gatos dentro de temas. Qual seria a ninhada na sua especialidade profissional?

Eu começo: Briefing, Job, Brainstorm e Learning
November 5, 2025 at 12:57 PM
AWS announces Spark Connect support on Amazon EMR on EC2, enabling interactive PySpark development from local IDEs or SageMaker Unified Studio while Spark runs on a dedicated cluster.
Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere
AWS announces Spark Connect support on Amazon EMR on EC2, enabling interactive PySpark development from local IDEs or SageMaker Unified Studio while Spark runs on a dedicated cluster.
aws-news.com
September 23, 2026 at 6:30 PM
I asked ChatGPT to fix some PySpark I thought I wrote in pseudocode. Imagine my horror when it replied, "Your syntax is already correct."
January 6, 2026 at 1:38 AM
After two years of full stack, I’ve moved into a data engineering role. Looking forward to learning everything PySpark 🐍 Maybe I’ll even resurrect my blog.
December 7, 2024 at 10:25 AM
PySpark for Beginners: Building Intermediate-Level Skills

towardsdatascience.com/pyspark-for-...
PySpark for Beginners: Building Intermediate-Level Skills | Towards Data Science
A practical next step into partitions, shuffles, joins, caching, and execution plans.
towardsdatascience.com
July 11, 2026 at 3:03 AM
Glue + PySpark + SQL - what could go wrong!?!?!

Unescaped single quotes is what.
March 4, 2026 at 7:54 PM
Looks like PySpark, Ray, and Dask have a competitor - Bodo. Has anyone tried it yet? This is the claimed execution time.
github.com/bodo-ai/Bodo
December 16, 2024 at 7:18 PM
Debug une Stacktrace Java provoquée par un code Python… telle est là vie que j'ai décidé de mener.

Merci PySpark.
January 22, 2026 at 4:33 PM
sparkはサーバー立てるのがクソ面倒かちゃ。でもpysparkはほぼpandasらしいかちゃ(pysparkつかっちゃことない)
January 9, 2025 at 12:13 PM