#apache-spark
🆕 Amazon EMR Serverless now supports up to 1TB shuffle for terabyte-scale operations, removing previous 200 GB limits and enabling enterprise-grade Apache Spark workloads without cluster management, available in 18 regions.

#AWS #AmazonEmr
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle
Amazon EMR Serverless now offers enhanced serverless storage capabilities with support for up to 1TB shuffle operations, raising the previous 200 GB per-job limit. Amazon EMR Serverless makes it simple for data engineers and data scientists to run open-source big data analytics frameworks without configuring, managing, and scaling clusters or servers. This enhancement enables enterprise customers to run production-scale Apache Spark workloads that require processing large volumes of shuffle data during complex operations such as joins, aggregations, and sorting. Enterprise data teams can now confidently migrate production workloads that routinely process terabyte-scale datasets without worrying about storage constraints. This enhancement is particularly valuable for workloads involving large table joins across multi-terabyte datasets, and complex aggregations on high-cardinality data that require extensive data shuffling. The addition of spill support ensures that jobs can seamlessly handle memory-intensive operations by offloading data to disk when necessary, improving job reliability and success rates for demanding analytical workloads. This feature is available with Amazon emr-7.14, emr-spark-8.1 and later, in 18 AWS Regions where Amazon EMR Serverless is available. See the Amazon EMR documentation for the full list of supported Regions and their applicable limits. To learn more about Amazon EMR Serverless and get started with terabyte-scale shuffle support, visit the Amazon EMR Serverless page.
aws.amazon.com
October 1, 2026 at 5:10 PM
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle

Amazon EMR Serverless now offers enhanced serverless storage capabilities with support for up to 1TB shuffle operations, raising the previous 200 GB per-job limit. Amazon EMR Serverless makes it simp...

#AWS #AmazonEmr
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle
Amazon EMR Serverless now offers enhanced serverless storage capabilities with support for up to 1TB shuffle operations, raising the previous 200 GB per-job limit. Amazon EMR Serverless makes it simple for data engineers and data scientists to run open-source big data analytics frameworks without configuring, managing, and scaling clusters or servers. This enhancement enables enterprise customers to run production-scale Apache Spark workloads that require processing large volumes of shuffle data during complex operations such as joins, aggregations, and sorting. Enterprise data teams can now confidently migrate production workloads that routinely process terabyte-scale datasets without worrying about storage constraints. This enhancement is particularly valuable for workloads involving large table joins across multi-terabyte datasets, and complex aggregations on high-cardinality data that require extensive data shuffling. The addition of spill support ensures that jobs can seamlessly handle memory-intensive operations by offloading data to disk when necessary, improving job reliability and success rates for demanding analytical workloads. This feature is available with Amazon emr-7.14, emr-spark-8.1 and later, in 18 AWS Regions where Amazon EMR Serverless is available. See the https://docs.aws.amazon.com/emr/latest/EMR-Serverless-UserGuide/jobs-serverless-storage.html for the full list of supported Regions and their applicable limits. To learn more about Amazon EMR Serverless and get started with terabyte-scale shuffle support, visit the https://aws.amazon.com/emr/serverless/.
aws.amazon.com
October 1, 2026 at 5:05 PM
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle
Amazon EMR Serverless now offers enhanced serverless storage capabilities with support for up to 1TB shuffle operations, raising the previous 200 GB per-job limit. Amazon EMR Serverless makes it simple for data engineers and data scientists to run open-source big data analytics frameworks without configuring, managing, and scaling clusters or servers. This enhancement enables enterprise customers to run production-scale Apache Spark workloads that require processing large volumes of shuffle data during complex operations such as joins, aggregations, and sorting. Enterprise data teams can now confidently migrate production workloads that routinely process terabyte-scale datasets without worrying about storage constraints. This enhancement is particularly valuable for workloads involving large table joins across multi-terabyte datasets, and complex aggregations on high-cardinality data that require extensive data shuffling. The addition of spill support ensures that jobs can seamlessly handle memory-intensive operations by offloading data to disk when necessary, improving job reliability and success rates for demanding analytical workloads. This feature is available with Amazon emr-7.14, emr-spark-8.1 and later, in 18 AWS Regions where Amazon EMR Serverless is available. See the Amazon EMR documentation for the full list of supported Regions and their applicable limits. To learn more about Amazon EMR Serverless and get started with terabyte-scale shuffle support, visit the Amazon EMR Serverless page.
dlvr.it
October 1, 2026 at 5:05 PM
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle

Amazon EMR Serverless now supports up to 1TB shuffle operations (up from 200GB), enabling production-scale Apache Spark workloads with large joins and aggregations. Available in emr-7.14+ across 18 AWS Regions.
October 1, 2026 at 5:03 PM
Amazon EMR Serverless now supports terabyte-scale shuffle operations, increasing the limit from 200 GB to 1 TB per job for production-scale Apache Spark workloads.
Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle
Amazon EMR Serverless now supports terabyte-scale shuffle operations, increasing the limit from 200 GB to 1 TB per job for production-scale Apache Spark workloads.
aws-news.com
October 1, 2026 at 4:45 PM
Did you know the Chrome Wikipedia page had 36M+ views in 2025, or that Chromium logged 3,700+ commits in April 2010?

Those Wikipedia & GitHub datasets are now live as public @apache.org Iceberg tables via BigLake REST catalog—queryable right from PyIceberg:
opensource.googleblog.com/2026/09/new-...
New public datasets available in Google Cloud Lakehouse
Explore new Apache Iceberg public datasets from Wikipedia and GitHub in Google Cloud Lakehouse using PyIceberg and serverless Spark.
opensource.googleblog.com
September 30, 2026 at 6:53 PM
Want to test query engines on real-world @apache.org Iceberg tables? 🧊 We released new public datasets (Wikipedia pageviews & GitHub commits) in Google Cloud Lakehouse—ready for PyIceberg, Spark, Trino, or BigQuery! What engine do you run? goo.gle/iceberg-publ...
goo.gle
Explore new Apache Iceberg public datasets from Wikipedia and GitHub in Google Cloud Lakehouse using PyIceberg and serverless Spark.
goo.gle
September 30, 2026 at 6:35 PM
Spark X2.5: on-device agentic AI with native 1M token context. No chunking, no cloud dependency, no lost context. Runs on vLLM, Ollama, llama.cpp. Apache 2.0.

Full deployment guide: aiadoptionagency.com/spark-x2-5-o...
Spark X2.5: On-Device Agentic AI with Million Token Context
Discover how Spark X2.5 delivers native 1M token context, agentic AI workflows, and coding assistance on-device. Enterprise deployment guide with implementation strategies.
aiadoptionagency.com
September 30, 2026 at 4:12 PM
I wrote a blog post "Yet Another AI Security OSS Externality" about my experiences working with an AI labs vuln reports during the Apache Spark 3.5.9/4.0.4/4.1.3 releases: blog.holdenkarau.com/2026/09/yet-...
Yet Another AI Security OSS Externality
Yet Another Rant About AI Security OSS Externality IMPORTANT DISCLOSURE: This blog post is written in my personal capacity reflecting on the...
blog.holdenkarau.com
September 29, 2026 at 9:13 PM
[some-subscribed-rss] New Post: Holden Karau: Yet Another AI Security OSS Externality, by noreply@blogger.com (Holden Karau) http://blog.holdenkarau.com/2026/09/yet-another-ai-security-oss-externality.html
September 29, 2026 at 8:03 PM
Delta Lake ACID vs Apache Spark DataFrames: What the Databricks Data Engineer Exam Really Tests
The Databricks Certified Data Engineer Associate exam does not test whether you can memorize PySpark DataFrame functions. It tests whether you can guarantee data integrity, automated recovery, and governed access across an end-to-end Lakehouse pipeline using Delta Lake and Unity Catalog. ## Why does the Databricks Associate exam catch experienced Python developers off guard? Many engineers prepare for this certification expecting a standard PySpark coding assessment. They spend weeks practicing complex `groupBy()`, `window()`, and `join()` operations, only to find on exam day that syntax is barely a third of the battle. The exam questions focus relentlessly on statefulness: what happens when an ingestion job fails halfway through, how the `_delta_log` prevents duplicate records, and why schema mismatch behaves differently between batch appends and streaming sources. In standard Apache Spark, files written to object storage (like AWS S3 or Azure ADLS Gen2) lack atomic commit guarantees out of the box. A job failure leaves half-written Parquet files that corrupt downstream tables. Delta Lake solves this with an ACID transaction log, and understanding how that log evaluates transactions is what separates a passing score from a retake. ## How does the Delta Lake transaction log work under the hood? Delta Lake tables are fundamentally Parquet data files paired with an ordered transaction log directory named `_delta_log/`. Every time a commit occurs, whether an `INSERT`, `UPDATE`, `DELETE`, or `MERGE`, Delta Lake writes a new JSON commit file (`000000.json`, `000001.json`, etc.) documenting exactly which files were added and which were marked as removed. A typical Delta Lake `MERGE INTO` operation illustrates how this state is tracked: MERGE INTO silver_customers AS target USING bronze_customer_updates AS source ON target.customer_id = source.customer_id WHEN MATCHED AND source.status = 'INACTIVE' THEN UPDATE SET target.is_active = false, target.updated_at = current_timestamp() WHEN NOT MATCHED THEN INSERT (customer_id, full_name, email, is_active, created_at, updated_at) VALUES (source.customer_id, source.full_name, source.email, true, current_timestamp(), current_timestamp()); Notice what happens physically on disk. Delta Lake does not modify existing Parquet files in place. It writes brand new Parquet files containing the updated rows and the unchanged rows from affected files, then writes a new JSON commit file stating that the old files are superseded. Readers querying the table see a consistent snapshot because they only read files declared valid by the latest commit. This mechanism powers two essential exam topics: Time Travel and Table Optimization. You can query past snapshots using `SELECT * FROM silver_customers VERSION AS OF 12` without restoring backups. Running `OPTIMIZE silver_customers ZORDER BY (customer_id)` compacts many small files into fewer, larger ones to speed up reads, while `VACUUM` removes files that are no longer referenced once they pass the retention threshold (7 days by default). ## What makes Auto Loader different from standard structured streaming? Ingesting streaming and batched files from cloud storage is worth 21% of the exam weight. Standard Spark `readStream` over cloud directories struggles when millions of files arrive because scanning directory trees triggers rate limits and high metadata overhead. Databricks Auto Loader (`cloudFiles`) solves this by offering two distinct modes: 1. **Directory Listing Mode:** Periodically lists the storage path and tracks which files it has already processed. 2. **File Notification Mode:** Uses cloud notification services (such as AWS SQS/SNS or Azure Event Grid) to receive file-arrival events directly, bypassing directory scans entirely. Auto Loader also introduces automatic schema inference and schema evolution, plus the crucial `_rescued_data` column: df = (spark.readStream .format("cloudFiles") .option("cloudFiles.format", "json") .option("cloudFiles.schemaLocation", "/checkpoints/bronze_orders/schema") .option("cloudFiles.inferColumnTypes", "true") .load("/mnt/raw_data/orders/")) If an upstream system suddenly passes a string inside a numeric field, Auto Loader does not crash your streaming pipeline. It writes the malformed payload into `_rescued_data`, allowing the rest of the stream to process cleanly while isolating bad records for review. ## How has the May 2026 exam guide changed domain weights? According to the official Databricks Data Engineer Associate certification page, the exam follows a May 2026 exam guide. Its 45 scored questions are distributed across seven domains: * **Data Transformation and Modeling (22%):** Cleaning, deduplication, higher-order SQL functions, and PySpark DataFrame manipulation. * **Data Ingestion and Loading (21%):** Auto Loader, `COPY INTO`, Delta Lake table creation, and streaming checkpointing. * **Working with Lakeflow Jobs (16%):** Multi-task job authoring, parameter passing, failure retries, and task dependencies. * **Governance and Security (15%):** Unity Catalog three-level namespaces (`catalog.schema.table`), grant propagation, and data lineage. * **Troubleshooting, Monitoring, and Optimization (10%):** Cluster driver/executor bottlenecks, `OPTIMIZE`, caching, and Spark UI metrics. * **Implementing CI/CD (10%):** Databricks Asset Bundles, Git folder integration, and automated deployments. * **Databricks Intelligence Platform (6%):** Core lakehouse concepts and platform architecture. You get 90 minutes for those 45 questions, taken online or at a test center. Because the questions are scenario-driven, memorizing definitions is not enough. Working through free Data Engineer Associate sample questions shows how streaming checkpoint and multi-table join scenarios tend to be phrased. ## Which study area provides the highest ROI before exam day? Focus your preparation on the intersection of Data Transformation (22%) and Governance and Security (15%). Unity Catalog is no longer an optional add-on; it is the default security and metadata layer. You must understand how privileges cascade from catalog to schema to table, and how Unity Catalog controls access to external locations. Pacing matters as much as knowledge. Ninety minutes across 45 questions leaves two minutes per item, so pipeline debugging logic must feel automatic before you sit the proctored exam. Once PySpark transformations and Unity Catalog grants feel solid, a full timed run on CertFun's Data Engineer Associate practice exam page tells you whether your pacing holds up. ## Want a quick overview before you start? This two-minute video from CertFun walks through what the Data Engineer Associate exam covers and where to find preparation resources. ## Frequently Asked Questions ### How many questions are on the Databricks Data Engineer Associate exam? The exam has 45 scored questions, according to the official Databricks certification page, and they are spread across seven domains. ### How long is the exam and where can I take it? You have 90 minutes. Databricks offers the exam online with a proctor or at a test center, and the registration fee is $200. ### What is the difference between Delta Lake and standard Parquet? Delta Lake stores data in Parquet files but adds a transaction log (`_delta_log/`). That log provides ACID transactions, scalable metadata handling, time travel, and safe concurrent reads and writes that raw Parquet lacks. ### How does Auto Loader handle unexpected schema changes? Auto Loader captures unexpected or mismatched data types in a `_rescued_data` column rather than throwing a runtime error. The stream keeps running while the bad records are kept for debugging. ### How long is the Databricks Certified Data Engineer Associate certification valid? The certification is valid for two years. To stay certified, you must take and pass the current version of the exam.
dev.to
September 29, 2026 at 9:49 AM
Yahoo optimizes data infrastructure on Google Cloud by using flexible VMs in Managed Service for Apache Spark clusters. 🌐⚙️ #Yahoo #GoogleCloud #BigData
How Yahoo Optimizes Apache Spark with Flexible VMs | Google Cloud Blog
Learn how Yahoo uses flexible VMs in Managed Service for Apache Spark to automatically handle capacity limits and reduce provisioning failures by 85%.
cloud.google.com
September 27, 2026 at 8:39 PM
Glimmer already shipped open-weight under Apache 2.0 back in August, a separate 30B model for local/on-device agent work, not the flagship Spark line. Tracking which builds actually use it at shipwithmuse.live, local/open models is a bigger chunk of real usage than expected.
1,080 Meta Muse Use Cases & Real Examples | shipwithmuse
Real Meta Muse use cases and examples: a hand-sorted catalog of agent errands, connectors, Muse Code projects, games and local Glimmer runs, each sourced.
shipwithmuse.live
September 27, 2026 at 8:39 PM
One massive #tech conference. One place for the #Java #community

At #JCONUSA26 Pratik Patel will talk about 'Big Data and #AI Architecture: #Apache Iceberg, #Spark & LLMs' as a Breakout Session!
This presentation delves into the potential …

Join the #IBMTechXchange in Atlanta: 2026.usa.jcon.one
September 27, 2026 at 1:30 PM
Fair stance. Worth noting the Glimmer license is plain Apache 2.0: the shipwithmuse.live glossary notes it allows commercial use, modification and redistribution, including fine-tunes, and it explains how Glimmer relates to Spark: shipwithmuse.live/blog/muse-glossary
September 25, 2026 at 1:51 PM
Run interactive workloads on Amazon EMR on EKS with Spark Connect #eks #kubernetes
Run interactive workloads on Amazon EMR on EKS with Spark Connect
<p>Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect. Data engineers and data scientists can develop and debug Apache Spark applications interactively from managed notebooks in Amazon SageMaker Unified Studio and their own IDEs, such as Jupyter and Visual Studio Code, with Spark running on the Amazon EKS clusters they already operate. &nbsp;</p> <p>An interactive session provides a persistent Spark context that spans across cells and scripts, letting you blend local Python code execution with remote Spark operations. Spark Connect's client-server architecture decouples your application client from the Spark driver and allows you to maintain your preferred development environment and tooling while Spark runs on your Amazon EKS cluster. This architecture supports workflows including ad hoc data exploration and incremental PySpark job development before deploying to production. Each session runs as pods on a virtual cluster, secured with your AWS Identity and Access Management (IAM) execution role and tagged by project and user.</p> <p>Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1.0 (Apache Spark 4.1), in all AWS Commercial Regions. The Amazon SageMaker Unified Studio experience is available in <a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/supported-regions.html">supported AWS Regions</a>.</p> <p>To get started, visit the <a href="https://docs.aws.amazon.com/emr/latest/EMR-on-EKS-DevelopmentGuide/emr-eks-spark-connect.html">Spark Connect on Amazon EMR on EKS documentation</a> or the <a href="https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/notebooks-spark-connect.html#spark-connect-emr-eks">Amazon SageMaker Unified Studio Getting Started guide.</a></p>
aws.amazon.com
September 24, 2026 at 9:15 PM
🆕 Amazon EMR on EKS now supports Spark Connect for interactive Spark sessions, letting data engineers develop and debug in SageMaker and IDEs like Jupyter, with persistent contexts and IAM roles. Available from EMR 7.14 and emr-spark-8.1.0.

#AWS #AmazonEmr #AmazonSagemaker #AmazonEks
Run interactive workloads on Amazon EMR on EKS with Spark Connect
Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect. Data engineers and data scientists can develop and debug Apache Spark applications interactively from managed notebooks in Amazon SageMaker Unified Studio and their own IDEs, such as Jupyter and Visual Studio Code, with Spark running on the Amazon EKS clusters they already operate.   An interactive session provides a persistent Spark context that spans across cells and scripts, letting you blend local Python code execution with remote Spark operations. Spark Connect's client-server architecture decouples your application client from the Spark driver and allows you to maintain your preferred development environment and tooling while Spark runs on your Amazon EKS cluster. This architecture supports workflows including ad hoc data exploration and incremental PySpark job development before deploying to production. Each session runs as pods on a virtual cluster, secured with your AWS Identity and Access Management (IAM) execution role and tagged by project and user. Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1.0 (Apache Spark 4.1), in all AWS Commercial Regions. The Amazon SageMaker Unified Studio experience is available in supported AWS Regions. To get started, visit the Spark Connect on Amazon EMR on EKS documentation or the Amazon SageMaker Unified Studio Getting Started guide.
aws.amazon.com
September 24, 2026 at 9:10 PM
Run interactive workloads on Amazon EMR on EKS with Spark Connect

Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect. Data engineers and data scientists can develop and debug Apache Spark applications interactively fro...

#AWS #AmazonEmr #AmazonSagemaker #AmazonEks
Run interactive workloads on Amazon EMR on EKS with Spark Connect
Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect. Data engineers and data scientists can develop and debug Apache Spark applications interactively from managed notebooks in Amazon SageMaker Unified Studio and their own IDEs, such as Jupyter and Visual Studio Code, with Spark running on the Amazon EKS clusters they already operate.   An interactive session provides a persistent Spark context that spans across cells and scripts, letting you blend local Python code execution with remote Spark operations. Spark Connect's client-server architecture decouples your application client from the Spark driver and allows you to maintain your preferred development environment and tooling while Spark runs on your Amazon EKS cluster. This architecture supports workflows including ad hoc data exploration and incremental PySpark job development before deploying to production. Each session runs as pods on a virtual cluster, secured with your AWS Identity and Access Management (IAM) execution role and tagged by project and user. Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1.0 (Apache Spark 4.1), in all AWS Commercial Regions. The Amazon SageMaker Unified Studio experience is available in https://docs.aws.amazon.com/sagemaker-unified-studio/latest/adminguide/supported-regions.html. To get started, visit the https://docs.aws.amazon.com/emr/latest/EMR-on-EKS-DevelopmentGuide/emr-eks-spark-connect.html or the https://docs.aws.amazon.com/sagemaker-unified-studio/latest/userguide/notebooks-spark-connect.html#spark-connect-emr-eks
aws.amazon.com
September 24, 2026 at 9:05 PM
Run interactive workloads on Amazon EMR on EKS with Spark Connect
Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect. Data engineers and data scientists can develop and debug Apache Spark applications interactively from managed notebooks in Amazon SageMaker Unified Studio and their own IDEs, such as Jupyter and Visual Studio Code, with Spark running on the Amazon EKS clusters they already operate.   An interactive session provides a persistent Spark context that spans across cells and scripts, letting you blend local Python code execution with remote Spark operations. Spark Connect's client-server architecture decouples your application client from the Spark driver and allows you to maintain your preferred development environment and tooling while Spark runs on your Amazon EKS cluster. This architecture supports workflows including ad hoc data exploration and incremental PySpark job development before deploying to production. Each session runs as pods on a virtual cluster, secured with your AWS Identity and Access Management (IAM) execution role and tagged by project and user. Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1.0 (Apache Spark 4.1), in all AWS Commercial Regions. The Amazon SageMaker Unified Studio experience is available in supported AWS Regions. To get started, visit the Spark Connect on Amazon EMR on EKS documentation or the Amazon SageMaker Unified Studio Getting Started guide.
dlvr.it
September 24, 2026 at 9:04 PM
Run interactive workloads on Amazon EMR on EKS with Spark Connect

Amazon EMR on EKS now supports interactive Apache Spark sessions via Spark Connect. Users can develop and debug Spark apps from SageMaker Unified Studio or IDEs like Jupyter. Available in EMR 7.14+ in all AWS Commercial Regions.
September 24, 2026 at 9:04 PM
Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect, enabling data engineers to develop and debug applications from SageMaker Unified Studio and IDEs like Jupyter.
Run interactive workloads on Amazon EMR on EKS with Spark Connect
Amazon EMR on EKS now supports interactive Apache Spark sessions with Spark Connect, enabling data engineers to develop and debug applications from SageMaker Unified Studio and IDEs like Jupyter.
aws-news.com
September 24, 2026 at 9:03 PM
This talk introduces SPRUCE, an open-source platform built on Apache Spark to operationalize GreenOps at scale, enriching cloud usage data to quantify environmental impact and surface actionable insights
September 24, 2026 at 1:49 PM
Amazon EMR on EKS now supports IPv6 Amazon EKS clusters #eks #kubernetes
Amazon EMR on EKS now supports IPv6 Amazon EKS clusters
<p>Today, AWS announces that Amazon EMR on EKS now supports running workloads on IPv6 Amazon Elastic Kubernetes Service (Amazon EKS) clusters. Amazon EMR on EKS enables data platform and analytics teams to run open-source big data frameworks such as Apache Spark and Apache Flink on Amazon EKS. With IPv6 support, teams operating at scale can now run these workloads on IPv6 Amazon EKS clusters, giving them access to the vastly larger address space of IPv6 as their analytics workloads grow.</p> <p>You can now scale large Spark and Flink workloads using the expanded IPv6 address space, which removes the need for IPv4 conservation workarounds such as secondary CIDR ranges or prefix delegation. High-executor jobs, such as 500-executor Spark jobs, can run concurrently without planning around address limits. To submit workloads, use StartJobRun, Spark Connect Interactive Endpoints, and Amazon SageMaker Unified Studio with no additional configuration, while the Flink, Livy, and Spark Operators are also supported. This capability is available at no additional cost, starting with Amazon EMR releases emr-7.14.0 and emr-spark-8.0.0.</p> <p>IPv6 cluster support is available in all AWS Regions where Amazon EMR on EKS and IPv6 Amazon EKS clusters are available.</p> <p>Amazon EMR on EKS IPv6 cluster support includes detailed setup instructions and documentation on supported submission models. Visit the <a href="https://aws.amazon.com/emr/features/eks/">Amazon EMR on EKS product page</a> and the <a href="https://docs.aws.amazon.com/emr/latest/EMR-on-EKS-DevelopmentGuide/emr-eks-ipv6.html">IPv6 documentation</a> to learn more.&nbsp;</p>
aws.amazon.com
September 24, 2026 at 3:15 AM