Data Engineer Associate Practice Questions
Prepare for DATA-ENG-ASSOC with more than an answer.
- Exam fee
- $200 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 5
- Databricks Intelligence Platform10%
- Development and Ingestion30%
- Data Processing & Transformations31%
- Productionizing Data Pipelines18%
- Data Governance & Quality11%
- 1
A data engineer is analyzing a slow-running Spark query using the Spark UI. They navigate to the Stages tab and observe a single, long-running stage with a very large number of tasks. The Directed Acyclic Graph (DAG) for the stage shows a wide transformation, such as a
groupByon a high-cardinality key. What is the most common performance bottleneck indicated by this pattern?Show answer details
Correct answer: B
Wide transformations, like
groupBy,join, orrepartition, require data to be redistributed or 'shuffled' among the worker nodes across the network. A single stage with many tasks performing such an operation indicates a large shuffle. This is a common and significant performance bottleneck in Spark because network I/O is much slower than memory access. The solution often involves optimizing the query to reduce the amount of data being shuffled. - 2
A data engineering team is setting up an Auto Loader pipeline to process text-based log files that arrive continuously. The logs have a consistent, well-defined CSV format. To optimize cost and processing latency, the team wants Auto Loader to use the most efficient mechanism for discovering new files in a cloud storage bucket that contains millions of existing files. Which
cloudFiles.modeshould they configure?Show answer details
Correct answer: C
File notification mode is the most scalable and efficient method for discovering new files in large storage locations. It subscribes to file event notifications from the cloud provider (e.g., S3 Event Notifications, Azure Event Grid). Instead of repeatedly listing all files in a directory (which is slow and expensive with millions of files), it processes new files as events are received. This significantly reduces latency and cost for continuous processing.
- 3
A data engineer needs to calculate the 7-day moving average of product sales from a
daily_salesDataFrame, which has columnssale_date,product_id, andsales_amount. The calculation must be performed independently for eachproduct_id. Which PySpark code snippet correctly computes this moving average?Show answer details
Correct answer: A
This is the correct approach. It uses a PySpark Window function, which is the standard way to perform calculations over a sliding range of rows.
partitionBy("product_id")ensures the calculation restarts for each product.orderBy("sale_date")sorts the data correctly.rowsBetween(-6, Window.currentRow)defines the 7-day window (the current row plus the previous 6 rows). Finally,avg("sales_amount").over(windowSpec)applies the average aggregation over this defined window. - 4
A project requires building and deploying a machine learning pipeline which involves several notebooks for feature engineering, model training, and batch inference. The entire set of artifacts, including notebooks, Python libraries, and the job definition, needs to be version-controlled in Git and deployed automatically across dev, staging, and prod environments. Which modern Databricks feature provides a declarative, bundled approach to manage and deploy this entire project as a single unit?
Show answer details
Correct answer: C
Databricks Asset Bundles (DAB) are the intended solution for this problem. They allow developers to define all project artifacts—notebooks, jobs, libraries, cluster definitions, and environment-specific configurations—in a declarative YAML file (
databricks.yml). This bundle can be checked into Git and deployed as a single, consistent unit using the Databricks CLI, which is ideal for CI/CD automation. - 5
What are the key benefits of using serverless compute for Databricks Workflows? (Select TWO)
Show answer details
Correct answer: B, C
This is a primary benefit. With serverless compute, Databricks manages the cluster sizing, optimization, and patching, providing a 'hands-off' experience that reduces operational overhead.
Serverless compute uses a warm pool of resources managed by Databricks, which significantly reduces the startup latency for jobs compared to traditional clusters that need to provision new VMs from the cloud provider for each run.
- 6
A data engineering team is migrating its development workflow from the Databricks UI to a local IDE using Databricks Connect. A junior engineer successfully sets up their connection profile but receives a
Py4JErrorupon trying to initialize aSparkSession. The cluster they are connecting to runs Databricks Runtime 14.3, which uses Python 3.11.2. The engineer's local environment is running Python 3.11.5. What is the primary reason for this connection failure?Show answer details
Correct answer: B
Databricks Connect requires that the major and minor Python versions of the local client environment match the version on the Databricks cluster exactly. In this case, the local version is 3.11.5 while the cluster version is 3.11.2. Even though the major version (3) and minor version (11) match, the strict requirement often extends to the patch version for full compatibility, but the major/minor mismatch is the key principle. An expired PAT would result in an authentication error, not a
Py4JError. A firewall issue would likely cause a timeout. An incorrect cluster ID would result in a 'not found' error. - 7
A DevOps team is implementing a CI/CD pipeline to deploy a multi-task Databricks workflow using Databricks Asset Bundles (DAB). The pipeline must handle deployments to development, staging, and production workspaces, each with different compute policies and secret scopes. Which TWO components of the
databricks.ymlfile are essential for managing these environment-specific configurations? (Select TWO)Show answer details
Correct answer: B, E
The
targetssection is specifically designed to manage environment-specific configurations. Each target can have its own workspace URL, root path, and variable definitions, allowing the same bundle to be deployed to different environments with the correct settings.Variables, defined within each target, allow for parameterization of the bundle's resources. This is how you would specify a different cluster policy ID for production versus development, or reference a different secret scope for database credentials in each environment.
- 8
A streaming pipeline using Auto Loader is configured to ingest JSON files from a cloud storage location. The pipeline runs successfully for several weeks but suddenly fails. Investigation of the
_rescuecolumn reveals that several recent files contain records where a previously numerictransaction_amountfield is now a string (e.g.,"100.50"instead of100.50). The desired behavior is to automatically adapt the target Delta table schema to accommodate this change without manual intervention. Which Auto Loader option should have been configured to handle this situation gracefully?Show answer details
Correct answer: D
The
rescuemode forcloudFiles.schemaEvolutionModeis designed for this exact scenario. It instructs Auto Loader to infer the schema and, if a data type change is detected (like numeric to string), it adds the new column to the_rescuecolumn. For schema evolution, the correct option is to enable schema evolution on the write stream using.option("mergeSchema", "true")and handle rescued data. However, among the given choices,rescuemode is the feature specifically designed to handle columns with mixed data types by rescuing them, which aligns with the observed behavior. WhilemergeSchemais also needed,schemaEvolutionModerescuedirectly addresses the problem of incompatible data types appearing in a column. - 9
A data architect is designing a Medallion architecture for a financial services company. The raw data (Bronze layer) contains sensitive personally identifiable information (PII). The Silver layer must contain the same records but with all PII columns pseudonymized. The Gold layer will contain aggregated data with no PII. The compliance team requires that only a specific service principal, used by an automated cleansing job, can read the raw PII data from the Bronze layer. All other users and groups should be denied access. How should this security requirement be implemented using Unity Catalog?
Show answer details
Correct answer: B
Unity Catalog operates on a default-deny model. To meet the requirement, you should avoid granting any broad permissions at the catalog or schema level. The correct approach is to grant the necessary, specific privileges (
SELECTto read,MODIFYmight be needed for the job's operations) directly to the service principal on the target Bronze tables. This ensures that only that principal can access the sensitive data, enforcing the principle of least privilege. - 10
True or False: When a Databricks job cluster is configured with a cluster pool, it can start faster because it acquires its driver and worker nodes from the pool of idle instances, reducing the time spent waiting for the cloud provider to provision new virtual machines.
Show answer details
Correct answer: A
This statement is true. The primary purpose of cluster pools is to reduce cluster start and auto-scaling times by maintaining a set of idle, ready-to-use instances. When a job requests a cluster attached to a pool, it gets its nodes from this warm pool instead of requesting new instances from the cloud provider, which significantly shortens the startup latency.
