Modern cloud data platforms rely heavily on Apache Spark for large-scale data processing. However, as data volumes grow, traditional Spark deployments face performance limitations due to the Java Virtual Machine (JVM). These limits include high CPU overhead from row-based processing, delays from Java garbage collection, and an inability to fully utilize modern hardware features such as Single Instruction, Multiple Data (SIMD) vectorization. To solve these problems, data engineers are turning to native acceleration layers. These engines bypass the JVM by rewriting core execution paths in low-level languages such as C++ and Rust, enabling faster, columnar processing directly on the hardware. The scope of this thesis is an evaluation of these native acceleration layers in cloud environments. Specifically, this research compares Databricks Photon, a proprietary C++ execution engine, with Apache DataFusion Comet, an open-source Rust-based plugin for Apache Spark. While Photon is a premium, vendor-locked accelerator, Comet offers a community-driven alternative that speeds up standard Spark deployments without tying users to a specific platform. To evaluate these technologies, this study uses a structured benchmarking methodology based on the industry-standard TPC-DS workload at a 100 GB Scale Factor. The experiments isolate and analyze query speed, horizontal scalability, and cost-efficiency across five different setups: standard open-source Spark, open-source Spark with Comet, Amazon EMR Spark, standard Databricks, and Databricks with Photon. By testing these setups across different cluster sizes (1-16 worker nodes) on AWS r5d.xlarge infrastructure, this thesis maps how each engine scales. Ultimately, this work provides a clear framework to help organizations choose between open-source flexibility and proprietary performance while managing their Total Cost of Ownership (TCO).

Modern cloud data platforms rely heavily on Apache Spark for large-scale data processing. However, as data volumes grow, traditional Spark deployments face performance limitations due to the Java Virtual Machine (JVM). These limits include high CPU overhead from row-based processing, delays from Java garbage collection, and an inability to fully utilize modern hardware features such as Single Instruction, Multiple Data (SIMD) vectorization. To solve these problems, data engineers are turning to native acceleration layers. These engines bypass the JVM by rewriting core execution paths in low-level languages such as C++ and Rust, enabling faster, columnar processing directly on the hardware. The scope of this thesis is an evaluation of these native acceleration layers in cloud environments. Specifically, this research compares Databricks Photon, a proprietary C++ execution engine, with Apache DataFusion Comet, an open-source Rust-based plugin for Apache Spark. While Photon is a premium, vendor-locked accelerator, Comet offers a community-driven alternative that speeds up standard Spark deployments without tying users to a specific platform. To evaluate these technologies, this study uses a structured benchmarking methodology based on the industry-standard TPC-DS workload at a 100 GB Scale Factor. The experiments isolate and analyze query speed, horizontal scalability, and cost-efficiency across five different setups: standard open-source Spark, open-source Spark with Comet, Amazon EMR Spark, standard Databricks, and Databricks with Photon. By testing these setups across different cluster sizes (1-16 worker nodes) on AWS r5d.xlarge infrastructure, this thesis maps how each engine scales. Ultimately, this work provides a clear framework to help organizations choose between open-source flexibility and proprietary performance while managing their Total Cost of Ownership (TCO).

Performance and Cost of Native Spark Accelerators: Photon vs. Comet

PIROLO, ALESSANDRO
2025/2026

Abstract

Modern cloud data platforms rely heavily on Apache Spark for large-scale data processing. However, as data volumes grow, traditional Spark deployments face performance limitations due to the Java Virtual Machine (JVM). These limits include high CPU overhead from row-based processing, delays from Java garbage collection, and an inability to fully utilize modern hardware features such as Single Instruction, Multiple Data (SIMD) vectorization. To solve these problems, data engineers are turning to native acceleration layers. These engines bypass the JVM by rewriting core execution paths in low-level languages such as C++ and Rust, enabling faster, columnar processing directly on the hardware. The scope of this thesis is an evaluation of these native acceleration layers in cloud environments. Specifically, this research compares Databricks Photon, a proprietary C++ execution engine, with Apache DataFusion Comet, an open-source Rust-based plugin for Apache Spark. While Photon is a premium, vendor-locked accelerator, Comet offers a community-driven alternative that speeds up standard Spark deployments without tying users to a specific platform. To evaluate these technologies, this study uses a structured benchmarking methodology based on the industry-standard TPC-DS workload at a 100 GB Scale Factor. The experiments isolate and analyze query speed, horizontal scalability, and cost-efficiency across five different setups: standard open-source Spark, open-source Spark with Comet, Amazon EMR Spark, standard Databricks, and Databricks with Photon. By testing these setups across different cluster sizes (1-16 worker nodes) on AWS r5d.xlarge infrastructure, this thesis maps how each engine scales. Ultimately, this work provides a clear framework to help organizations choose between open-source flexibility and proprietary performance while managing their Total Cost of Ownership (TCO).
2025
Performance and Cost of Native Spark Accelerators: Photon vs. Comet
Modern cloud data platforms rely heavily on Apache Spark for large-scale data processing. However, as data volumes grow, traditional Spark deployments face performance limitations due to the Java Virtual Machine (JVM). These limits include high CPU overhead from row-based processing, delays from Java garbage collection, and an inability to fully utilize modern hardware features such as Single Instruction, Multiple Data (SIMD) vectorization. To solve these problems, data engineers are turning to native acceleration layers. These engines bypass the JVM by rewriting core execution paths in low-level languages such as C++ and Rust, enabling faster, columnar processing directly on the hardware. The scope of this thesis is an evaluation of these native acceleration layers in cloud environments. Specifically, this research compares Databricks Photon, a proprietary C++ execution engine, with Apache DataFusion Comet, an open-source Rust-based plugin for Apache Spark. While Photon is a premium, vendor-locked accelerator, Comet offers a community-driven alternative that speeds up standard Spark deployments without tying users to a specific platform. To evaluate these technologies, this study uses a structured benchmarking methodology based on the industry-standard TPC-DS workload at a 100 GB Scale Factor. The experiments isolate and analyze query speed, horizontal scalability, and cost-efficiency across five different setups: standard open-source Spark, open-source Spark with Comet, Amazon EMR Spark, standard Databricks, and Databricks with Photon. By testing these setups across different cluster sizes (1-16 worker nodes) on AWS r5d.xlarge infrastructure, this thesis maps how each engine scales. Ultimately, this work provides a clear framework to help organizations choose between open-source flexibility and proprietary performance while managing their Total Cost of Ownership (TCO).
Apache Spark
Native Acceleration
TPC-DS Benchmark
File in questo prodotto:
File Dimensione Formato  
AlessandroPirolothesis.pdf

Accesso riservato

Dimensione 3.35 MB
Formato Adobe PDF
3.35 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110962