Apache-Spark-Developer Associate Developer for Apache Spark Practice Questions
Prepare for Apache-Spark-Developer with more than an answer.
- Exam fee
- $200 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 7
- Apache Spark Architecture and Components20%
- Using Spark SQL20%
- Developing Apache Spark DataFrame/DataSet API Applications30%
- Troubleshooting and Tuning Apache Spark DataFrame API Applications10%
- Structured Streaming10%
- Using Spark Connect to deploy applications5%
- Using Pandas API on Apache Spark5%
- 1
What is the key advantage of using the Pandas API on Spark compared to running pandas on a single machine?
Show answer details
Correct answer: B
The primary advantage of the Pandas API on Spark is that it provides a familiar pandas-like API that operates on Spark DataFrames. This allows data scientists and engineers to use their existing pandas knowledge to process datasets that are too large to fit in the memory of a single machine by distributing the computation across a Spark cluster.
- 2
Which two statements accurately describe lazy evaluation in Apache Spark? (Select TWO)
Show answer details
Correct answer: B, C
- 3
True or False: In Structured Streaming, using
trigger(once=True)creates a batch job that processes all available new data since the last trigger and then stops, effectively turning a streaming query into a triggerable, incremental batch job.Show answer details
Correct answer: A
This statement is true. The
Trigger.Once(ortrigger(once=True)) setting configures the streaming query to run as a single micro-batch. It will identify and process all new data that has arrived since the last time it ran, update the state and sink, and then shut down the query. This is a common pattern for converting a streaming job into an incremental batch job that can be scheduled by an external orchestrator like Databricks Jobs or Airflow. - 4
An organization is migrating its Spark workloads to a new environment. They have a variety of applications, including interactive data analysis from notebooks, automated batch jobs, and remote applications running in different IDEs. They want a unified way for all these clients to connect to a single, managed Spark cluster. Which deployment mode or technology is best suited for this requirement?
Show answer details
Correct answer: C
Spark Connect is designed specifically for this use case. It provides a decoupled client-server architecture, allowing various clients (IDEs, notebooks, applications) to connect remotely to a Spark cluster through a stable API. This centralizes cluster management and allows developers to work from their preferred environments without needing to configure complex cluster access on each client machine.
- 5
A data engineer is working with a DataFrame
dfthat has duplicate rows. The engineer wants to create a new DataFrame that contains only the unique rows fromdf. Which of the following operations should be used?Show answer details
Correct answer: C
The
distinct()method returns a new DataFrame containing only the unique rows from the original DataFrame. It is an alias fordropDuplicates()when no columns are specified. This operation is a wide transformation as it requires a shuffle to identify and remove duplicate rows across all partitions. - 6
A data engineering team is developing a Spark application to process sensitive financial data. They need to pass a large, read-only lookup table (approximately 500MB) containing currency exchange rates to all executor nodes. This table is used in a join operation within multiple tasks. Which Spark feature should be used to distribute this lookup table efficiently and minimize network I/O?
Show answer details
Correct answer: B
A broadcast variable is the correct choice for distributing a large, read-only dataset to all worker nodes. The driver serializes the variable and sends it to each executor only once, where it is cached in memory. Tasks on that executor can then access the data locally without causing repeated network transfers. Accumulators are used for aggregating results back to the driver. Caching the DataFrame would work, but broadcast is specifically designed for this 'side data' distribution pattern and is generally more efficient for joins. A UDF is for custom logic, not data distribution.
- 7
In the context of the Apache Spark execution hierarchy, which of the following events will always trigger the creation of a new Spark Stage?
Show answer details
Correct answer: C
Spark creates a new stage at each shuffle boundary. Wide transformations, such as
groupBy(),join(), orrepartition(), require data to be redistributed (shuffled) across the network between executors. This redistribution marks the end of one stage and the beginning of another. Narrow transformations can be pipelined within a single stage. Actions trigger the execution of a job, which consists of one or more stages, but the action itself doesn't define the stage boundary; the shuffle does. - 8
A Spark job processing a large dataset is experiencing performance degradation. Analysis of the Spark UI shows that one task in a particular stage is taking significantly longer than all other tasks. The stage involves a
groupBy('user_id')operation. What is the most likely cause of this issue and the most appropriate solution? (Select TWO)Show answer details
Correct answer: B, D
- 9
A developer needs to read a large Parquet dataset partitioned by
year,month, andday. To optimize read performance, they only want to load data for the first week of January 2023. Which Spark SQL query correctly applies partition pruning to achieve this?SELECT * FROM sales_parquet WHERE _______Show answer details
Correct answer: B
For partition pruning to be effective, the filter conditions must be applied directly to the partition columns. This allows Spark's Catalyst optimizer to read the file system metadata and skip reading the data files in partitions that do not match the filter. Applying functions to the partition columns (like
concatorto_date) can prevent the optimizer from pushing down the predicate, forcing a full table scan. The correct approach is to filter directly on the raw partition columns. - 10
A streaming application needs to calculate a running count of events per user and output the updated count for each user as new data arrives in every micro-batch. Which Structured Streaming output mode is designed for this use case?
Show answer details
Correct answer: C
Update mode is specifically designed for use cases involving aggregations. In this mode, only the rows that were updated in the result table since the last trigger will be written to the sink. This is perfect for outputting running counts where only the users with new events in the current micro-batch will have their counts updated in the output. Complete mode would rewrite the entire result table, and Append mode is not supported for aggregations without a watermark.
