Spark writing to S3 failed: java.lang.NoSuchMethodError:...
Symptom:When using Spark writing to S3, the insert query failed:Caused by: java.lang.NoSuchMethodError:...
View ArticleSpark writing to S3 failed: java.lang.NoSuchMethodError:...
Symptom:When using Spark writing to S3, the insert query failed:java.lang.NoSuchMethodError:...
View ArticleHow to access Azure Open Dataset from Spark
Goal:This article explains how to access Azure Open Dataset from Spark. Env:spark-3.1.1-bin-hadoop2.7Solution:Microsoft Azure Open Dataset is curated and cleansed data - including weather, census, and...
View ArticleUnderstand Decimal precision and scale calculation in Spark using GPU or CPU...
Goal:This article research on how Spark calculates the Decimal precision and scale using GPU or CPU mode. Basically we will test Addition/Subtraction/Multiplication/Division/Modulo/Union in this...
View Articlekubelet failed to start after rebooting
Symptom:kubelet failed to start after rebooting. Env:Ubuntu 18.04Kubernetes 1.19 Root Cause:From "journalctl -xefu kubelet", we can find out the root cause:kubelet[11111]: F0430 xx:xx:xx.123456 11111...
View ArticleHow to use Spark Operator to run Spark job with Rapids Accelerator
Goal:This article shares the steps on how to run Spark job with Rapids Accelerator using Spark Operator in a Kubernetes Cluster.Env:Spark 3.1.1Rapids Accelerator 0.4.1 with cuDF 0.18.1Kubernetes...
View ArticleRapids Accelerator compatibility related to...
Goal:This article talked about the compatibility of Rapids Accelerator for Spark regarding parquet writing related to parameters spark.sql.legacy.parquet.datetimeRebaseModeInWrite etc.Env:Spark...
View ArticleSpark Code -- Dig into SparkListenerEvent
Goal:This article digs into different types of SparkListenerEvent in Spark event log with some examples. Understanding this can help us know how to pares Spark event log.Env:Apache Spark 3.1.1 source...
View ArticleHow to use latest version of Rapids Accelerator for Spark on EMR
Goal:This article shows how to use latest version of Rapids Accelerator for Spark on EMR. Currently the latest EMR 6.2 only ships with Rapids Accelerator 0.2.0 with cuDF 0.15 jar.However as of today,...
View ArticleHow to use NVIDIA Nsight Systems to profile a Spark on K8s job with Rapids...
Goal:This article explains how to use NVIDIA Nsight Systems to profile a Spark on K8s job with Rapids Accelerator.This is a follow-up blog after How to use NVIDIA Nsight Systems to profile a Spark job...
View ArticleHow to use NVIDIA Nsight Systems to profile a Spark job on Rapids Accelerator
Goal:This article explains how to use NVIDIA Nsight Systems to profile a Spark job on Rapids Accelerator.Env:Spark 3.1.1 (Standalone Cluster)RAPIDS Accelerator for Apache Spark 0.5 snapshotcuDF jar...
View ArticleHow to enable GpuKryoRegistrator on RAPIDS Accelerator for Spark
Goal:This article shares the steps to enable GpuKryoRegistrator on RAPIDS Accelerator for Spark.Env:Spark 3.1.1RAPIDS Accelerator for Apache Spark 0.4.1Solution:As mentioned in Spark Tuning Doc:Java...
View ArticleHow to install a Kubernetes Cluster with NVIDIA GPU on AWS using DeepOps
Goal:This article shares a step-by-step guide on how to install a Kubernetes Cluster with NVIDIA GPU on AWS using DeepOps.Env:AWS EC2 (G4dn)Ubuntu 18.04Solution: Most of the steps are the same as...
View ArticleHow to install a Kubernetes Cluster with NVIDIA GPU on AWS
Goal:This article shares a step-by-step guide on how to install a Kubernetes Cluster with NVIDIA GPU on AWS. It includes spinning up an AWS EC2 instance, installing NVIDIA drivers&cudatoolkit,...
View Articleconcat_ws example on Spark with RAPIDS Accelerator
Goal:This is a quick example of operator contact_ws on Spark with RAPIDS Accelerator.Env:Spark 3.1.1RAPIDS Accelerator for Apache Spark 0.4.1Solution:1. concat_ws can convert an Array of Strings to a...
View ArticleHands-on native cuDF Pandas UDF
Goal:This article will help show some hands-on steps to play with native cuDF Pandas UDF on Spark with RAPIDS Accelerator for Apache Spark.Env:RAPIDS Accelerator for Apache Spark 0.4.1Spark 3.1.1RTX...
View ArticleHow to run the pandas cudf_udf test for RAPIDS Accelerator for Apache Spark
Goal:How to run the pandas cudf_udf test for RAPIDS Accelerator for Apache Spark.Env:RAPIDS Accelerator for Apache Spark 0.4Spark 3.1.1Solution: 1. Compile RAPIDS Accelerator for Apache Spark1.a Create...
View ArticleUnderstanding RAPIDS Accelerator For Apache Spark's supported timezone
Goal:This article explains the current supported timezone for "RAPIDS Accelerator For Apache Spark".Env:RAPIDS Accelerator For Apache Spark 0.4Concept:As per current 0.4 Doc mentions: operations...
View ArticleSpark Tuning -- Adaptive Query Execution(3): Dynamically optimizing skew joins
Goal:This article explains Adaptive Query Execution (AQE)'s "Dynamically optimizing skew joins" feature introduced in Spark 3.0. This is a follow up article for Spark Tuning -- Adaptive Query...
View ArticleWhat Dataset API is not supported for RAPIDS Accelerator for Apache Spark
Goal:This article explains what Dataset API is not supported for RAPIDS Accelerator for Apache Spark.Env:Spark 3.0.2RAPIDS Accelerator for Apache Spark 0.3Solution: Currently RAPIDS Accelerator for...
View Article