PMI-CPMAI Practice Questions
Prepare for PMI-CPMAI with more than an answer.
- 1
Which of the following statements best describes the relationship between Deep Learning and Machine Learning?
Show answer details
Correct answer: C
The correct hierarchy is that Artificial Intelligence (AI) is the broad field, Machine Learning (ML) is a subset of AI, and Deep Learning (DL) is a further subset of ML. Deep Learning utilizes multi-layered artificial neural networks (hence 'deep') to learn from vast amounts of data, enabling it to solve more complex problems than some traditional ML algorithms.
- 2
A project is using a supervised learning model to classify customer support tickets into categories like 'Billing', 'Technical Issue', and 'Feedback'. The dataset used for training is known to have a severe class imbalance, with 'Billing' tickets making up 80% of the data. During Phase V (Model Evaluation), which metric would be the most misleading indicator of the model's performance on the minority classes?
Show answer details
Correct answer: A
Accuracy is the most misleading metric in cases of severe class imbalance. A naive model that simply classifies every ticket as 'Billing' would achieve 80% accuracy but would be completely useless for identifying 'Technical Issue' or 'Feedback' tickets. Metrics like Precision, Recall, and especially the F1-score (or a macro-averaged score) provide a much better assessment of performance across all classes, including the rare ones.
- 3
True or False: The CPMAI methodology is a rigid, waterfall-style process where each of the six phases must be fully completed and signed off before the next one can begin.
Show answer details
Correct answer: B
This statement is false. A core principle of the CPMAI methodology is that it is highly iterative and agile. While it provides a structured six-phase framework, it is expected that project teams will loop back to earlier phases as they learn more. For example, insights gained during Model Development (Phase IV) might require revisiting Data Preparation (Phase III) or even Business Understanding (Phase I).
- 4
What is the primary difference between supervised and unsupervised machine learning?
Show answer details
Correct answer: B
The fundamental distinction lies in the data used for training. Supervised learning algorithms learn from data that has been manually labeled with the correct outcomes or targets (e.g., images of cats labeled 'cat'). The goal is to learn a mapping function to predict the output for new, unseen data. Unsupervised learning algorithms, in contrast, work with unlabeled data and try to find inherent patterns or structures within it, such as grouping similar data points together (clustering).
- 5
A project team is using a third-party API for sentiment analysis as part of a larger application. The project manager needs to assess the risk of 'model robustness' related to this component. Which scenario best illustrates a failure in model robustness?
Show answer details
Correct answer: C
Model robustness refers to a model's ability to maintain its performance level when faced with small, unexpected, or adversarial perturbations in the input data. In this case, a minor change (a common misspelling and extra punctuation) causes the model's prediction to flip from positive to negative. This indicates a lack of robustness. API downtime is a reliability issue, and cost increase is a procurement issue, not issues of model robustness.
- 6
A utility company is developing an AI model to predict power outages based on weather data and sensor readings from the grid. In which phase of the CPMAI lifecycle would the team perform Exploratory Data Analysis (EDA) to identify correlations, anomalies, and initial patterns in the historical data?
Show answer details
Correct answer: B
Phase II: Data Understanding is dedicated to the initial collection and exploration of data. Activities like Exploratory Data Analysis (EDA), data quality assessment, and initial pattern discovery are central to this phase. The goal is to gain familiarity with the data and identify potential challenges and opportunities before proceeding to intensive data preparation and modeling.
- 7
A data scientist has trained two models for a binary classification task. To compare their performance irrespective of the classification threshold, they have plotted the following ROC curves. Based on the diagram, what can the project manager conclude?
graph TD subgraph ROC Curve Analysis A[Model A (AUC = 0.92)] B[Model B (AUC = 0.78)] C(Random Classifier (AUC = 0.50)) endShow answer details
Correct answer: B
The Area Under the ROC Curve (AUC) represents a model's ability to discriminate between classes. A value of 1.0 indicates a perfect classifier, while 0.5 indicates a classifier with no discriminative ability (equivalent to random guessing). Since Model A has a higher AUC (0.92) than Model B (0.78), it has a superior overall performance in separating the classes across all possible thresholds.
- 8
A financial services firm is in Phase IV (Model Development) of a CPMAI project to create a real-time fraud detection system. The data science team has developed a highly accurate deep learning model. However, during a review, the compliance team raises a concern that the model's decisions are completely opaque, violating new regulatory requirements for 'Right to Explanation'. What is the most appropriate next step for the project manager according to the CPMAI methodology?
Show answer details
Correct answer: C
The CPMAI methodology is iterative. A critical new requirement, such as regulatory compliance for explainability, discovered in a later phase necessitates an iteration. The correct action is to loop back within the current phase or a previous one to address the gap. In this case, returning to Model Development (Phase IV) to build a compliant model is the right approach. Ignoring the requirement (A) or trying to change the project's fundamental goals (B) is inappropriate. Scrapping the model entirely (D) is too drastic; iteration is preferred.
- 9
An agricultural AI project aims to predict crop yield based on satellite imagery, weather patterns, and soil sensor data. The dataset is characterized by high dimensionality, non-linear relationships, and significant interaction between features. The project sponsor requires a model that is both highly accurate and provides clear insights into which factors are most influential on the yield. Which algorithm would be the most suitable choice?
Show answer details
Correct answer: D
Gradient Boosted Trees, particularly implementations like XGBoost, are well-suited for this problem. They excel at handling complex, non-linear data with high dimensionality and feature interactions, typically yielding high accuracy. Crucially, they also provide built-in feature importance metrics, which directly addresses the sponsor's requirement for insights. Linear Regression assumes linear relationships, which is not the case here. K-Means is an unsupervised clustering algorithm, unsuitable for this supervised prediction task. SVM can handle non-linearity but is less interpretable than tree-based methods.
- 10
A project manager is overseeing the development of a data pipeline for a large-scale AI system that will process both real-time streaming data from IoT devices and nightly batch data from a legacy CRM. Key requirements are scalability, fault tolerance, and the ability to manage complex data workflows. Which combination of technologies is most appropriate for this use case? (Select TWO)
Show answer details
Correct answer: A, C
Apache Kafka is a distributed streaming platform designed to handle high-throughput, real-time data feeds with excellent fault tolerance, making it ideal for ingesting IoT data.
Apache Airflow is a workflow orchestration tool that allows for programmatic authoring, scheduling, and monitoring of complex data pipelines, including both batch and streaming jobs. It satisfies the need to manage complex workflows.
