Skip to content

ML-ASSOC Machine Learning Associate Practice Questions

Prepare for ML-ASSOC with more than an answer.

210 questions in the full set20 sample questionsUpdated Jan 29, 2026
Exam fee
$200 USD
Level
Associate
Valid for
2 years
Domains covered on the exam 4
  1. Databricks Machine Learning38%
  2. Data Processing19%
  3. Model Development31%
  4. Model Deployment12%
  1. 1

    In a Spark ML Pipeline, what is the fundamental difference between an Estimator and a Transformer?

    Show answer details

    Correct answer: B

    This is the core distinction. An Estimator (e.g., LogisticRegression, StringIndexer) has a .fit() method that takes a DataFrame and learns from it to produce a model, which is a Transformer. A Transformer (e.g., a trained model, StandardScalerModel, OneHotEncoder) has a .transform() method that takes a DataFrame and returns a new DataFrame with a column appended or changed.

  2. 2

    True or False: Databricks Model Serving endpoints can only serve models that have been logged to MLflow using one of its built-in model flavors, such as mlflow.pyfunc.

    Show answer details

    Correct answer: A

    This statement is true. Databricks Model Serving relies on the pyfunc model flavor as a standardized interface for loading models and serving predictions. This allows the serving infrastructure to handle any model type (scikit-learn, TensorFlow, custom Python models, etc.) in a generic way, as long as it's wrapped in the pyfunc flavor.

  3. 3

    A data scientist has just completed a classification model training run. In addition to logging parameters and metrics, they want to save a PNG image of the model's confusion matrix to their MLflow run for visual inspection. Which MLflow Client API function should be used for this purpose?

    Show answer details

    Correct answer: C

    mlflow.log_artifact() is the correct function for logging arbitrary files, such as images, text files, or model plots, to an MLflow run. log_metric is for numerical values, log_param is for key-value parameters, and log_model is for saving the model itself in a specific format.

  4. 4

    While k-fold cross-validation provides a more robust estimate of model performance than a simple train-validation split, it is not always the best choice. What are two potential disadvantages of using k-fold cross-validation? (Select TWO)

    Show answer details

    Correct answer: A, C

    Because the model must be trained k times (once for each fold), the overall training time is roughly k times longer than a single train-test split.

    For time-series data, random shuffling of data into folds (as done in standard k-fold) violates the temporal order. This means the model could be trained on future data to predict the past, leading to an overly optimistic performance estimate. Specialized techniques like TimeSeriesSplit are required.

  5. 5

    For which type of machine learning algorithm can applying one-hot encoding to a high-cardinality categorical feature be particularly problematic due to the risk of multicollinearity if not handled correctly (e.g., by dropping one of the new columns)?

    Show answer details

    Correct answer: B

    Linear models, such as Linear and Logistic Regression, are sensitive to multicollinearity. When a categorical feature is one-hot encoded, the resulting binary columns are perfectly correlated (e.g., if a value is not in categories A, B, or C, it must be in D). This perfect correlation, known as the dummy variable trap, can make the model's coefficient estimates unstable and difficult to interpret. Dropping one of the encoded columns is a common way to resolve this.

  6. 6

    A large financial services company operates multiple Databricks workspaces for different business units. They need to develop a centralized repository of customer features (e.g., credit score, transaction frequency) that can be securely shared and reused across all workspaces, with strict access controls managed by a central governance team. Which approach best meets these requirements?

    Show answer details

    Correct answer: C

    Creating Feature Store tables within Unity Catalog at the account level is the correct approach. Unity Catalog provides a centralized governance and access control layer that spans across all workspaces within an account. This allows the central governance team to manage permissions on a single set of feature tables, which can then be securely accessed by authorized users and services from any workspace, fulfilling the core requirements of centralization and secure sharing.

  7. 7

    True or False: In the context of the bias-variance tradeoff, increasing a model's complexity (e.g., adding more layers to a neural network or increasing the depth of a decision tree) will generally decrease its bias but increase its variance.

    Show answer details

    Correct answer: A

    This statement is true. Increasing model complexity allows the model to learn more intricate patterns from the training data, which reduces its bias (the error from erroneous assumptions in the learning algorithm). However, a more complex model is also more likely to learn the noise in the training data, making it more sensitive to variations in the data. This increased sensitivity is known as higher variance, which can lead to overfitting and poor generalization to new, unseen data.

  8. 8

    A data scientist is cleaning a dataset containing employee salary information for a large corporation. They observe that the 'salary' column has a number of extreme outliers, including several C-level executive salaries that are orders of magnitude higher than the rest of the employees. For a feature engineering task, they need to impute a few missing salary values. Which imputation method should they prefer and why?

    Show answer details

    Correct answer: B

    The median is the correct choice because it is a robust measure of central tendency, meaning it is not significantly affected by extreme outliers. The mean, on the other hand, is sensitive to outliers; the very high executive salaries would pull the mean upwards, making it an unrepresentative value for the typical employee. Using the median ensures the imputed value is closer to the center of the majority of the data points.

  9. 9

    A leading e-commerce company wants to deploy a new product recommendation system. The system has several complex requirements:

    1. Real-time Personalization: The model must provide recommendations within 200ms of a user's action (e.g., viewing a product).
    2. Dynamic User Profiles: User feature vectors, which are used for inference, must be updated in near real-time based on their clickstream data.
    3. A/B Testing: The MLOps team must be able to deploy a new 'challenger' recommendation algorithm and route 10% of live traffic to it for evaluation against the current 'champion' model.
    4. Scalability: The system must handle traffic spikes during holiday seasons, scaling automatically without manual intervention.

    Which Databricks deployment architecture best fulfills all these requirements?

    graph TD subgraph User Interaction WebApp[Web Application] --> API_GW[API Gateway] end subgraph Real-time Processing Clickstream[Kafka: User Events] --> DLT[Delta Live Tables: User Profile Update] DLT --> OnlineStore[Online Feature Store] end subgraph Model Serving API_GW --> Endpoint[Model Serving Endpoint] Endpoint -- 90% --> Champion[Champion Model] Endpoint -- 10% --> Challenger[Challenger Model] Champion --> OnlineStore Challenger --> OnlineStore end subgraph Batch Processing ProductCatalog[Batch: Product Catalog] --> OfflineStore[Offline Feature Store] end

    Show answer details

    Correct answer: C

    This architecture correctly addresses all requirements. A serverless Model Serving endpoint provides low-latency (<200ms) inference and automatic scaling. Deploying multiple models to a single endpoint with traffic splitting directly enables A/B testing. An Online Feature Store, updated by a streaming process like DLT, provides the dynamic, low-latency feature vectors needed for real-time personalization.

  10. 10

    A machine learning engineer is using the FeatureEngineeringClient to create a new feature table in Unity Catalog. The table will store user features, and it's critical that each user is uniquely identified and that features can be looked up efficiently for online serving. Which parameter in the fe.create_table method is used to specify the unique identifier column(s) for the entities in the table?

    Show answer details

    Correct answer: B

    The primary_keys parameter is used to specify the column or columns that uniquely identify each row or entity in the feature table. This is crucial for the Feature Store as it uses these keys for joining features during training set creation and for efficient lookups in online stores.

Create an account to continue.