Continuing our "not an island" mood: stream remote datasets (HDF5, Zarr, B2Z) over HTTP/S3 lazily, and compute across formats in one expression: b2z + zarr * h5! Plus tiered caching & Win ARM64 wheels.
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
Continuing our "not an island" mood: stream remote datasets (HDF5, Zarr, B2Z) over HTTP/S3 lazily, and compute across formats in one expression: b2z + zarr * h5! Plus tiered caching & Win ARM64 wheels.
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
One lazy API for unite all remote data: browse B2Z/Zarr/HDF5 hierarchies over S3/GCS/HTTP, slice only what you need, cache in RAM or on disk, and export portable snapshots.
b2view demo below👇
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
One lazy API for unite all remote data: browse B2Z/Zarr/HDF5 hierarchies over S3/GCS/HTTP, slice only what you need, cache in RAM or on disk, and export portable snapshots.
b2view demo below👇
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
Remote arrays work over S3/GCS/HTTP via fsspec. lazy=True fetches only the blocks needed and reuses a validated cache — 5–17x faster in S3 benchmarks.
Also: concurrent Caterva2 writers + UTF-8 indexes.
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
Remote arrays work over S3/GCS/HTTP via fsspec. lazy=True fetches only the blocks needed and reuses a validated cache — 5–17x faster in S3 benchmarks.
Also: concurrent Caterva2 writers + UTF-8 indexes.
github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
On 1 Mrow with 100 distinct titles: group_by 22x faster, isin() 5.7x, 5.6x smaller on disk.
blosc.org/python-blosc2/guides/optimization_tips.html#repeated-text-store-it-as-dictionary
Enjoy data 🚀
On 1 Mrow with 100 distinct titles: group_by 22x faster, isin() 5.7x, 5.6x smaller on disk.
blosc.org/python-blosc2/guides/optimization_tips.html#repeated-text-store-it-as-dictionary
Enjoy data 🚀
New CTable nullability on a validity mask instead of a sentinel value ✨️
Also, wheels are now abi3 compliant w/ CPython 3.11+, including 3.15 (tested) 🚅
More info: github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
New CTable nullability on a validity mask instead of a sentinel value ✨️
Also, wheels are now abi3 compliant w/ CPython 3.11+, including 3.15 (tested) 🚅
More info: github.com/Blosc/python...
Compress Better, Compute Bigger 🚀
A Mandelbrot escape loop written plainly: ~155x faster than plain Python, ~2.3x than hand-vectorized NumPy.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #Blosc2
A Mandelbrot escape loop written plainly: ~155x faster than plain Python, ~2.3x than hand-vectorized NumPy.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #Blosc2
blosc2.field(blosc2.utf8())
string(200) costs 800 B/row on every read; utf8() rows cost what they weigh: ~4× less memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #DataScience #Blosc2
blosc2.field(blosc2.utf8())
string(200) costs 800 B/row on every read; utf8() rows cost what they weigh: ~4× less memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #DataScience #Blosc2
(rows + cols).compute(urlpath="big.b2nd", mode="w")
Chunk by chunk: similar speed, ~26× less peak memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #Blosc2
(rows + cols).compute(urlpath="big.b2nd", mode="w")
Chunk by chunk: similar speed, ~26× less peak memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #Blosc2
>>> arr = blosc2.open(path, mmap_mode="r")
Maps once, reads pages directly. ~15% faster for scattered reads, same memory. Free win, and scales with concurrent readers 🚀
blosc.org/python-blosc...
#Python #DataScience #Blosc2
>>> arr = blosc2.open(path, mmap_mode="r")
Maps once, reads pages directly. ~15% faster for scattered reads, same memory. Free win, and scales with concurrent readers 🚀
blosc.org/python-blosc...
#Python #DataScience #Blosc2
t["temperature"].sum(where=t.region == 3)
~1.9× faster, ~7× less memory than mask-then-index.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #DataScience #Blosc2
t["temperature"].sum(where=t.region == 3)
~1.9× faster, ~7× less memory than mask-then-index.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
#Python #DataScience #Blosc2
t["val"].sum() reduces the compressed column chunk-wise, never materialising it: 1.7× faster, ~12× less peak memory on 50M rows.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
t["val"].sum() reduces the compressed column chunk-wise, never materialising it: 1.7× faster, ~12× less peak memory on 50M rows.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
A @blosc2.dsl_kernel handed to lazyudf() compiles to native code, filling an NDArray chunk by chunk, in parallel. ~3.6x faster, ~2x less peak memory on 200M elements.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
A @blosc2.dsl_kernel handed to lazyudf() compiles to native code, filling an NDArray chunk by chunk, in parallel. ~3.6x faster, ~2x less peak memory on 200M elements.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
CTable.sort_by(view=True)[:k] and NDArray.iter_sorted(start=-k) read just that slice from the index sidecar instead of sorting everything: up to ~74× less time, ~193× less peak memory.
blosc.org/python-blosc...
Enjoy data!
#Python
CTable.sort_by(view=True)[:k] and NDArray.iter_sorted(start=-k) read just that slice from the index sidecar instead of sorting everything: up to ~74× less time, ~193× less peak memory.
blosc.org/python-blosc...
Enjoy data!
#Python
CTable auto-builds per-block min/max indexes for its scalar columns, so Column.min()/max() never decompress the column: ~4× faster, essentially no extra memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
CTable auto-builds per-block min/max indexes for its scalar columns, so Column.min()/max() never decompress the column: ~4× faster, essentially no extra memory.
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
A chunk-aligned read decompresses 1 chunk instead of 2 (~2.2× faster), and chunk-aligned slice() copies chunks as-is with no decompression at all (~4.9× faster). Same principle at block level.
blosc.org/python-blosc...
Enjoy data!
A chunk-aligned read decompresses 1 chunk instead of 2 (~2.2× faster), and chunk-aligned slice() copies chunks as-is with no decompression at all (~4.9× faster). Same principle at block level.
blosc.org/python-blosc...
Enjoy data!
Open with mmap_mode="r" in every reader — all share one set of mapped pages instead of a syscall + copy per read.
8 concurrent readers: ~4.5x faster, half the CPU, still ~one copy of the file in RAM.
blosc.org/python-blosc...
Enjoy data! 🚀
Open with mmap_mode="r" in every reader — all share one set of mapped pages instead of a syscall + copy per read.
8 concurrent readers: ~4.5x faster, half the CPU, still ~one copy of the file in RAM.
blosc.org/python-blosc...
Enjoy data! 🚀
blosc2.linspace(0, 1, N) fills a compressed array chunk by chunk — at 200M float64, ~25x less peak memory than asarray(np.linspace(...)), comparable speed.
Same for arange() and fromiter().
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
blosc2.linspace(0, 1, N) fills a compressed array chunk by chunk — at 200M float64, ~25x less peak memory than asarray(np.linspace(...)), comparable speed.
Same for arange() and fromiter().
blosc.org/python-blosc...
Compress Better, Compute Bigger 🚀
SWMR readers follow a writer's appends in near real-time, and opt-in locking makes updates atomic — no torn reads, no retries.
🎬 One writer appending, three readers chasing it live:
www.blosc.org/python-blosc...
#Python
SWMR readers follow a writer's appends in near real-time, and opt-in locking makes updates atomic — no torn reads, no retries.
🎬 One writer appending, three readers chasing it live:
www.blosc.org/python-blosc...
#Python
Newton-Raphson runs 2.7× faster on the browser! 💨
With #Pyodide, it’s not just Python in the browser: it’s the PyData stack 📊 too, #Blosc2 tacking compression 📦 + computing 🧮 duties.
Demo 👇
blosc.org/demos/newton...
Newton-Raphson runs 2.7× faster on the browser! 💨
With #Pyodide, it’s not just Python in the browser: it’s the PyData stack 📊 too, #Blosc2 tacking compression 📦 + computing 🧮 duties.
Demo 👇
blosc.org/demos/newton...
This is the new b2view Text User Interface that comes with latest Python-Blosc2 release.
github.com/Blosc/python...
Enjoy!
This is the new b2view Text User Interface that comes with latest Python-Blosc2 release.
github.com/Blosc/python...
Enjoy!
Under the hood we did amazing improvements, making toying with very large tabular data a pleasant, NumPy-centric experience.
Have fun!
Under the hood we did amazing improvements, making toying with very large tabular data a pleasant, NumPy-centric experience.
Have fun!
🖥️ b2view: interactive TUI browser
🎯 SUMMARY indexes → fast WHERE
🧮 DSL kernels as CTable columns
⚡ Faster chunk-by-chunk writes
🔀 where() via miniexpr
Compressed, NumPy-native arrays & tables.
📝 github.com/Blosc/python...
Enjoy data!
#Python #NumPy
🖥️ b2view: interactive TUI browser
🎯 SUMMARY indexes → fast WHERE
🧮 DSL kernels as CTable columns
⚡ Faster chunk-by-chunk writes
🔀 where() via miniexpr
Compressed, NumPy-native arrays & tables.
📝 github.com/Blosc/python...
Enjoy data!
#Python #NumPy
New in this release: sparse coordinate getters for faster random access, header-only b2nd metalayer access for plugins (no libblosc2 link needed), and new codec IDs for J2K/HTJ2K support. 🧩⚡
Release notes: github.com/Blosc/c-blos...
Compress better, compute bigger! 💥
As a result, much larger tables can be indexed. Look at how this fares against other good indexing engines in plots below.
Enjoy!
#TabularData
As a result, much larger tables can be indexed. Look at how this fares against other good indexing engines in plots below.
Enjoy!
#TabularData
Meet `CTable`: a new compressed columnar container with null support, fast queries, and easy Arrow/Parquet interoperability.
Compress better, analyze faster.
Release notes:
github.com/Blosc/python...
Meet `CTable`: a new compressed columnar container with null support, fast queries, and easy Arrow/Parquet interoperability.
Compress better, analyze faster.
Release notes:
github.com/Blosc/python...