AI Testing Practice Questions
Prepare for CT-AI with more than an answer.
- Level
- Specialist
- Valid for
- Lifetime (no expiration)
- 1
A project team is developing an AI system and is following the ML workflow. The team has just completed the 'Train the Model' step. According to the standard ML workflow, what are the next TWO immediate steps, performed iteratively, before the model is ready for independent testing against the holdout test dataset? (Select TWO)
flowchart LR A[Data Prep] --> B[Train Model] B --> C{...} C --> D{...} D --> B C --> E[Test Model]Show answer details
Correct answer: A, C
After the initial training, the model's performance must be evaluated using the validation dataset to see how well it generalizes. This step provides the metrics needed for the next step.
Based on the results from the evaluation step, the model's hyperparameters are adjusted (tuned) to improve performance. This evaluation and tuning process is often repeated multiple times before a final model candidate is chosen for testing.
- 2
A QA team is testing an AI-based system that uses a pre-trained model from a third-party vendor. The documentation for the pre-trained model is minimal. Which quality characteristic is most directly and negatively impacted by this lack of documentation, making it difficult for the QA team to assess risks like inherited bias or vulnerabilities?
Show answer details
Correct answer: C
Transparency refers to the ease with which the internal workings of the AI system, including the algorithm and training data, can be determined. Minimal documentation for a pre-trained model means there is low transparency. Without knowing what data the model was trained on or its architecture, it is extremely difficult for testers to assess potential inherited biases, understand its limitations, or anticipate its failure modes.
- 3
Which of the following describes a key difference between AI-based systems and conventional systems from a testing perspective?
Show answer details
Correct answer: B
This is the fundamental difference. In a conventional system, behavior is defined by explicit rules and logic (e.g., if-then-else statements) written by a programmer. In an AI-based system, particularly one using machine learning, the system learns its own rules and patterns from a dataset. This emergent behavior is not explicitly coded, which leads to unique testing challenges like the test oracle problem and the need to test the data itself.
- 4
A team is testing an AI-powered chatbot. They notice that when users ask the same question multiple times with slightly different phrasing, the chatbot sometimes provides inconsistent or contradictory answers. This makes it difficult to write traditional, assertion-based automated tests. The test lead suggests a technique where they compare the chatbot's response to the response from a different, established chatbot service for the same input questions. What is this technique called?
Show answer details
Correct answer: B
Back-to-back testing, also known as differential testing, involves running the same test inputs on two or more variants of a system and comparing the outputs. Using an established, trusted system as a reference (or pseudo-oracle) to compare against the system under test is a common implementation of this technique, especially for systems like chatbots where defining a single 'correct' answer is difficult.
- 5
The F1-score is often used as a performance metric for classification models. It is calculated as the harmonic mean of two other fundamental metrics. What are these two metrics?
Show answer details
Correct answer: C
The F1-score is defined as the harmonic mean of Precision and Recall. The formula is F1 = 2 * (Precision * Recall) / (Precision + Recall). It is particularly useful when you need to find a balance between Precision and Recall, especially when dealing with imbalanced datasets where accuracy can be misleading.
- 6
A financial institution has developed an AI model to assess loan application risk. The model is a deep neural network, making it a "black box" where the logic for a specific decision is not easily explainable. The test team faces a significant test oracle problem, as calculating the "correct" risk score for a new applicant profile is infeasible. They decide to use Metamorphic Testing (MT). Which of the following represents the MOST effective Metamorphic Relation (MR) for testing the logical consistency of this loan risk model?
Show answer details
Correct answer: C
This is the most effective Metamorphic Relation because it tests a fundamental logical property of a loan risk model without needing a precise expected output. The property is that, all other factors being equal, a higher income should not increase financial risk. This allows the team to verify the model's logical consistency even when the exact risk score calculation is unknown. The other options are less effective: a tiny income change may not trigger a score change, duplicating the test case only checks for determinism, and assuming proportionality is not a guaranteed property.
- 7
A QA team is testing a new AI-powered hiring tool that screens resumes to identify top candidates. The company is concerned about introducing unintentional bias against protected groups. The test data includes demographic information, but this data is NOT used as a feature for the model's prediction. Despite demographic data not being an input feature, the model may still exhibit bias. Which TWO of the following testing activities are most crucial for uncovering this hidden bias? (Select TWO)
Show answer details
Correct answer: B, C
This method, known as disparate impact analysis or statistical parity testing, is a primary technique for detecting outcome-based bias, regardless of the input features. It directly measures whether the model's decisions are fair in practice.
This is a critical step in identifying the root cause of hidden bias. Proxy variables are features that are not explicitly sensitive but act as stand-ins for protected characteristics, introducing bias through the training data itself.
- 8
A DevOps team wants to improve its testing of a complex REST API with hundreds of endpoints and intricate dependencies. Manual test case creation is slow and often misses complex interaction bugs. They decide to use an AI-based tool for test case generation that uses a search-based algorithm (like a genetic algorithm) to explore the API's behavior. What is the primary challenge the team will face when integrating this AI-driven test generation tool into their CI/CD pipeline?
Show answer details
Correct answer: C
This is the fundamental limitation of most AI-based test generation tools. While excellent at exploring the state space and generating novel inputs (the 'how to test'), they lack the domain-specific knowledge to determine what the correct output should be. This is known as the test oracle problem. The most common use for such tools is to find crashes or server errors (5xx), for which the oracle is simple, rather than to find subtle logical bugs.
- 9
A hospital is deploying an AI system to predict the likelihood of sepsis in ICU patients. Due to the critical nature of the decisions, regulations require that the model's predictions be explainable to clinicians. The development team chose a deep neural network (DNN) for its high accuracy. Which technique would be most appropriate for the testing team to use to validate the explainability requirement for this high-stakes, black-box model?
Show answer details
Correct answer: B
For a complex black-box model like a DNN, direct inspection of weights is not human-interpretable. LIME is designed specifically for this use case: providing local, understandable explanations for individual predictions of any black-box model. This allows a clinician to ask 'Why did the model flag this specific patient?' and get an answer based on the most influential features (e.g., 'because of high heart rate and low blood pressure'), making it ideal for validating explainability in a clinical setting.
- 10
True or False: Achieving 100% neuron coverage in a deep neural network guarantees that all logical paths within the model have been tested and that the model is free from defects.
Show answer details
Correct answer: B
The statement is false. Neuron coverage is a very weak structural coverage criterion, analogous to statement coverage in traditional code. It only ensures that each neuron has produced an output above a certain threshold at least once. It does not test the complex interactions between neurons, the different activation ranges, or the combinatorial logic of the network. Therefore, it provides no guarantee that the model is free of logical flaws or defects.
