Skip to content

ML-PRO Machine Learning Professional Practice Questions

Prepare for ML-PRO with more than an answer.

205 questions in the full set20 sample questionsUpdated Jan 29, 2026
Exam fee
$200 USD
Level
Professional
Valid for
2 years
Domains covered on the exam 3
  1. Model Development35%
  2. MLOps35%
  3. Model Deployment30%
  1. 1

    True or False: When a custom model is logged to Unity Catalog, its dependencies listed in the conda.yaml file are automatically registered as UC Volumes for improved governance.

    Show answer details

    Correct answer: B

    False. Unity Catalog governs data and AI assets like models, tables, and files (in Volumes). It does not manage Python package dependencies. The conda.yaml file is packaged as part of the model artifact itself, and the dependencies are installed from public or private repositories into the execution environment (like a Model Serving container) when the model is loaded, but they are not stored or governed as individual UC assets.

  2. 2

    Case Study:

    A media streaming company, Streamly, is enhancing its recommendation engine. They need to serve personalized recommendations in real-time with very low latency. The feature engineering process is complex, involving both historical user data (e.g., watch history, genre preferences) and real-time contextual data (e.g., time of day, device type).

    The historical features are large and pre-calculated nightly, stored in a Delta table. The real-time features are derived from the incoming user request. The ML team wants a unified solution to combine these two types of features at inference time. The solution must provide sub-100ms latency for feature lookups and serving.

    The team has evaluated several architectural patterns. They want to avoid managing separate infrastructure for online serving and ensure that the features used for training are consistent with those used for inference to prevent skew.

    Which Databricks architecture provides the most efficient and integrated solution for this hybrid feature-serving scenario?

    Show answer details

    Correct answer: A

    This architecture is the ideal solution. Databricks Online Tables are purpose-built for low-latency lookups of pre-computed features. Databricks Feature Serving is designed for on-demand computation of real-time features. The two services are designed to work together seamlessly within a Model Serving endpoint, providing a unified, high-performance solution for combining batch and real-time features at inference time.

  3. 3

    An ML engineer is building a monitor for an inference table using Lakehouse Monitoring. The goal is to evaluate model performance trends over time. Which of the following are required columns in the inference table for this type of analysis? (Select ALL that apply)

    Show answer details

    Correct answer: A, B, C, D

  4. 4

    A data science team is using MLflow to track a complex hyperparameter tuning experiment. They want to structure their logging so that there is one main parent run for the entire experiment, and each individual hyperparameter combination is logged as a nested child run. Which of the following code snippets correctly creates this nested run structure?

    Show answer details

    Correct answer: C

    To create nested runs, you must have an active parent run. The with mlflow.start_run(): statement creates the parent run. Inside this block, each subsequent call to with mlflow.start_run(nested=True): will create a new child run that is logically grouped under the parent in the MLflow UI.

  5. 5

    A team is setting up a multi-workspace MLOps architecture for development, staging, and production. To maintain consistency and control, they need to ensure that only models approved in the staging workspace's Unity Catalog can be promoted to the production catalog. Which Databricks feature is fundamental to enforcing this governance model?

    graph TD subgraph Dev Workspace A[Notebooks & Code] --> B(MLflow Tracking) B --> C{Dev UC Catalog} end subgraph Staging Workspace D[CI/CD: Run Tests] --> E(MLflow Tracking) E --> F{Staging UC Catalog} end subgraph Prod Workspace G[CI/CD: Deploy] --> H{Prod UC Catalog} H --> I[Model Serving] end C -- Promote --> F F -- Promote --> H

    Show answer details

    Correct answer: C

    Unity Catalog provides a centralized governance layer for all data and AI assets across multiple workspaces. By defining a single metastore for all workspaces, a CI/CD pipeline (typically running as a service principal) can be granted permissions to read models from the staging catalog and write/update them in the production catalog. This allows for a governed, auditable promotion process that is independent of individual user permissions and workspaces, which is the core of this architecture.

  6. 6

    A machine learning team is developing a model to predict customer churn. They are using Databricks Asset Bundles (DABs) to manage their project environments. They need to define separate configurations for development, staging, and production, including different cluster policies and secrets scopes. Which section of the databricks.yml file is specifically designed to manage these environment-specific overrides?

    Show answer details

    Correct answer: B

    The targets block in a databricks.yml file is used to define different deployment targets, which typically correspond to environments like development, staging, and production. This section allows for overriding default configurations specified in the main bundle or resources sections, enabling environment-specific settings for compute, secrets, and other parameters.

  7. 7

    An MLOps engineer is implementing a canary deployment for a new version of a demand forecasting model using Databricks Model Serving. The goal is to route 10% of the inference traffic to the new model version (version 2) while the remaining 90% goes to the stable version (version 1). Which configuration snippet correctly implements this traffic split within a model serving endpoint definition?

    Show answer details

    Correct answer: A

    Databricks Model Serving endpoints support traffic splitting for canary and blue-green deployments. The served_models array in the endpoint configuration is where you define which model versions are active. The traffic_percentage key for each model version entry specifies the percentage of requests that should be routed to it. This option correctly assigns 90% to version 1 and 10% to version 2.

  8. 8

    A data scientist is building a SparkML pipeline to process text data for sentiment analysis. The pipeline needs to tokenize text, remove stop words, and then convert the tokens into numerical feature vectors using TF-IDF. Which sequence of SparkML transformers is correct for this task?

    Show answer details

    Correct answer: C

    The correct logical sequence for this NLP preprocessing task is to first break the text into tokens (Tokenizer), then remove common stop words from the token list (StopWordsRemover), then convert the cleaned tokens into term frequencies (HashingTF), and finally re-weight the term frequencies based on their importance across the corpus (IDF).

  9. 9

    A team is building an automated retraining pipeline for a credit risk model. The pipeline should trigger a new training job whenever significant drift is detected in the model's key input features. They are using Lakehouse Monitoring to track drift. Which of the following components are essential for implementing this automated retraining workflow? (Select THREE)

    Show answer details

    Correct answer: A, B, C

  10. 10

    An ML engineer is tasked with creating a custom PyFunc model in MLflow. This model needs to load a pre-trained tokenizer from Hugging Face and a custom-trained scikit-learn classifier. The entire model, including the tokenizer, must be packaged together for deployment to a sandboxed environment without internet access. Which MLflow feature should be used to package the tokenizer along with the model?

    Show answer details

    Correct answer: C

    The artifacts parameter in mlflow.pyfunc.log_model() is designed for this exact purpose. It allows you to specify a dictionary of local file paths that will be packaged with the model. Inside the custom model's load_context method, you can then access these artifacts using the provided context object, ensuring the model is self-contained and portable.

Create an account to continue.