#apacheparquet
The ALP encoding is now officially part of @ApacheParquet.

parquet.apache.org/blog/2026/09...
September 23, 2026 at 11:06 AM
Great news for Arrow (and more work for us 😄).
Also, those work items also implicitly apply to #ApacheParquet. @julien.ledem.net
October 7, 2025 at 12:35 PM
Variant is coming soon to @ApacheParquet in Rust . Huge thanks to @db.cs.cmu.edu for getting the process started in @ApacheArrow with a great draft PR to kick off Variant support: github.com/apache/arrow... 🙏🙏🙏 Thank you
April 16, 2025 at 1:40 PM
Woot, woot: 2^8 🌟s for the #Hardwood GitHub repo 🥳! Might be a small step, but so great to see the interest in the project. #ApacheParquet

github.com/hardwood-hq/...
May 5, 2026 at 8:14 AM
Recording of "Introduction to Variant in @ApacheParquet ": www.youtube.com/watch?v=nlOJ...

Here are the slides: docs.google.com/presentation...
September 6, 2025 at 9:54 AM
You can use ApacheParquet for Vector Search with embedded indexes:

> We don’t change the file format; we just tune it.

Xiangpeng Hao explains how in blog.xiangpeng.systems/posts/vector...
February 10, 2026 at 12:17 PM
Check out @andrewlamb1111.bsky.social 's talk at the recent Iceberg meetup for a condensed overview of the the new Variant type coming to Parquet #apacheparquet #apachearrow
Recording of "Introduction to Variant in @ApacheParquet ": www.youtube.com/watch?v=nlOJ...

Here are the slides: docs.google.com/presentation...
September 7, 2025 at 6:19 PM
"From the One Billion Row Challenge to a Multi-Threaded Parquet Library"-an airhacks.fm podcast conversation with Gunnar Morling is ready to listen:
adambien.blog/roller/from_...
#java #podcast #apacheparquet #graalvm #airhacks
airhacks.fm podcast
podcast with adam bien
airhacks.fm
September 14, 2026 at 4:39 PM
#Hardwood is now #opensource.

A high-performance JVM library for reading #ApacheParquet, built with a multi-threaded architecture and zero mandatory external dependencies - a simpler alternative to the Apache Parquet Java implementation.

More details on #InfoQ 👉 bit.ly/3R8sviH

#Java #BigData
July 9, 2026 at 10:37 AM
Prateek Gaur and co at Snowflake reproduced the (great) results for the ALP encoding algorithm from CWI / Azim Afroozeh / Peter Boncz

ALP achieves ZSTD levels of compression and much faster decode. We are discussing adding it to @ApacheParquet: lists.apache.org/thread/tjtln...
October 17, 2025 at 1:05 PM
Thanks to @clflushopt.bsky.social, make massive TPCH datasets with tpchgen-cli 2.0:

SF1000 (1TB raw, 220GB in @ApacheParquet ) in less than 10 mins (6m45s) on aging laptop

Try it now:

pip install tpchgen-cli
tpchgen-cli --scale-factor 1000 --parts 100 --format=parquet

github.com/clflushopt/t...
September 4, 2025 at 12:51 PM
CRITICAL: Apache Parquet Java vulnerability (CVE-2025-46762) allows RCE; upgrade to 1.15.2 immediately. #ApacheParquet #RCE #Cybersecurity
Apache Parquet Java Flaw Exposes Systems To RCE
CRITICAL: Apache Parquet Java vulnerability (CVE-2025-46762) allows RCE; upgrade to 1.15.2 immediately. #ApacheParquet #RCE #Cybersecurity
securityonline.info
May 8, 2025 at 3:50 AM
1 Billion Row Challenge creator @gunnarmorling.dev is back with #Hardwood - a zero-dependency, ultra-fast #Java parser for #ApacheParquet.

🎧 Listen to the #InfoQ #podcast to learn more about Hardwood’s architecture, parallelization & performance optimization ⇨ bit.ly/4wP8jlR

#AI
Chasing Efficient Java Development: From 1BRC to Developing Hardwood AI Natively
Gunnar Morling, technologist at Confluent and Java Champion, shares his experiences with building high-performance applications in Java, especially in the data space. He shares insights from experimen...
bit.ly
May 26, 2026 at 10:55 AM
At @quantstack.bsky.social we designed novel bit-unpacking SIMD optimizations for @arrow.apache.org and #ApacheParquet, and implemented them entirely using C++ metaprogramming instead of Python-based code generation.

We'll publish a deep dive blog post soon.

github.com/apache/arrow...
GH-48277: [C++][Parquet] unpack with shuffle algorithm by AntoinePrv · Pull Request #47994 · apache/arrow
Rationale for this change The current bit-unpacking algorithm (which is implemented as a C++ code generator script in Python) does not fully leverage SIMD operations: all loads and some bitshifts u...
github.com
February 26, 2026 at 5:34 PM
I have so much work to do today but what I really want to do is kick the tires on the new DuckDB-Grafana plugin: github.com/motherduckdb.... This unlocks some really cool use cases involving Parquet and Arrow data. #DuckDB #ApacheArrow #ApacheParquet
GitHub - motherduckdb/grafana-duckdb-datasource
Contribute to motherduckdb/grafana-duckdb-datasource development by creating an account on GitHub.
github.com
January 24, 2025 at 6:11 PM
Centralized storage, decentralized compute: maximizing data use (OLAP) while minimizing ETL and disparate versions of data.

#hotTake #dataEngineering #dataAnalytics #dataScience #machineLearning #ai #cloudComputing #s3 #apacheArrow #apacheParquet #apacheIceberg
August 28, 2025 at 5:57 PM
Announcing the first ever @arrow.apache.org and #ApacheParquet meetup in Paris, kindly hosted by @datadoghq.com.

If you’re using Arrow or Parquet, looking for insights, or wanting to meet other community members, this meetup is for you. Please register if you plan to attend!

luma.com/6ed1oko1
Apache Arrow / Parquet - June 2026 meetup in Paris · Luma
Details We’re excited to announce the first ever Apache Arrow and Parquet meetup in Paris! This meetup will be hosted on June 18th by Datadog, in their…
luma.com
May 6, 2026 at 8:00 AM
🎤 Speaker Spotlight: Aditya Gadkari

We're excited to welcome Aditya Gadkari to #CascadiaRConf 2026!

Waiting on large datasets to load?
Aditya will demo how tools like Parquet can dramatically improve storage, read times, and workflows in R.
#rstats #DataScience #ApacheParquet #DataEngineering
June 16, 2026 at 3:31 PM
I am speaking at the #ApacheIceberg NYC Meetup on July 10th about Variant in
#ApacheParquet
which enable more efficient of processing semi structured data such as that found in JSON.

lu.ma/95a5qys1
NYC Apache Iceberg™ Community Meetup · Luma
🧊 Apache Iceberg Meetup is coming to the Big Apple! 🗽 Join us in NYC for an afternoon of ideas, innovation, and Iceberg. Whether you're building lakehouses…
lu.ma
July 3, 2025 at 9:43 AM
PyArrow 21 was a great release, especially for @hf.co users: PyArrow now seamlessly handles hf:// URIs and does content-defined chunking to reduce transfer and storage costs on HF. Check out this blog post: huggingface.co/blog/parquet... #apachearrow #apacheparquet
Parquet Content-Defined Chunking
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
August 5, 2025 at 3:49 PM
Every #TimeSeriesDatabase is a set of storage decisions: row layout, compression timing, partitioning - driving cost & query performance.

This #InfoQ article breaks down these fundamentals from first principles using #PostgreSQL & #ApacheParquet ⇨ bit.ly/42z4jbi

#TimeSeriesData #Database
May 14, 2026 at 4:14 AM
On April 1st, 2025, CVE-2025-30065 in Apache Parquet’s module was disclosed, with a CVSS score of 10.

Despite fears, exploitation is challenging with limited risk.

Discover what we found: https://go.f5.net/6vqac3fy

#F5Labs #ApacheParquet #Cybersecurity
June 2, 2025 at 6:00 PM
Apache Parquet and Java Spring Boot 👇

youtu.be/2rGtJTA4G38

If you're building data-heavy applications or optimizing backend performance, this one breaks down practical ways to integrate efficient storage formats with Spring Boot.

#ApacheParquet #SpringBoot #Java #Backend #DataEngineering
Spring Boot and Apache Parquet Make large data files easy to handle
YouTube video by Mike Møller Nielsen
youtu.be
March 30, 2026 at 9:00 AM