Skip to content

CompTIA Data+ (V2) Practice Questions

Prepare for DA0-002 with more than an answer.

159 questions in the full set12 sample questionsUpdated Jul 16, 2026

Unlock the full exam and previous versions

  • v1CompTIA Data+ (V2) 159 questions Current
  • DA0-001Legacy CompTIA Data+ 249 questions Locked
Exam fee
$255 USD
Time limit
90 minutes
Questions on the exam
Up to 90
Passing score
720 (scale 100-900)
Level
Entry-to-Mid Level
Valid for
3 years
Domains covered on the exam 5
  1. Data Concepts and Environments20%
  2. Data Acquisition and Preparation22%
  3. Data Analysis24%
  4. Visualization and Reporting20%
  5. Data Governance14%
  1. 1

    During data cleansing, an analyst discovers that a free-text comment field in a CRM database frequently contains unstructured US phone numbers formatted in various ways (e.g., 555-123-4567, (555) 1234567, 555.123.4567). The analyst needs to extract these phone numbers into a clean, standardized column. Which string manipulation technique is MOST appropriate for this task?

    Show answer details

    Correct answer: B

    Regular Expressions (RegEx) are designed for advanced pattern matching and extraction within strings. Because the phone numbers appear in various formats within unstructured text, RegEx is the only viable method to identify and extract the numerical patterns. Concatenation combines strings. Binning groups continuous values into categories. The TRIM function only removes leading or trailing whitespace.

  2. 2

    A data science team is evaluating a new machine learning algorithm designed to detect fraudulent credit card transactions. The null hypothesis (H0) states that a transaction is NOT fraudulent. In this specific context, what represents a Type II error, and why is it problematic?

    Show answer details

    Correct answer: B

    A Type II error occurs when you fail to reject a false null hypothesis (a False Negative). Since the null hypothesis is that the transaction is NOT fraudulent, a Type II error means the algorithm incorrectly accepts this and lets a fraudulent transaction pass. A Type I error (False Positive) would be rejecting the true null hypothesis (flagging a normal transaction as fraud).

    stateDiagram-v2 [*] --> Actual_Condition Actual_Condition --> Legitimate(H0_True) Actual_Condition --> Fraudulent(H0_False) Legitimate(H0_True) --> Flagged_Fraud(Reject_H0): Type I Error (False Positive) Legitimate(H0_True) --> Approved(Accept_H0): Correct Decision Fraudulent(H0_False) --> Approved(Accept_H0): Type II Error (False Negative) Fraudulent(H0_False) --> Flagged_Fraud(Reject_H0): Correct Decision
  3. 3

    A global logistics company uses an advanced analytics system. The system not only predicts which delivery trucks are likely to break down in the next week, but it also automatically schedules maintenance appointments and reroutes packages to other vehicles to minimize delivery delays. Which type of analysis is this system performing?

    Show answer details

    Correct answer: D

    Prescriptive analysis goes beyond predicting what will happen (predictive analysis) by recommending or automating actions to take advantage of the prediction. In this scenario, scheduling maintenance and rerouting packages are prescribed actions. Descriptive analysis explains what happened in the past. Diagnostic analysis explains why it happened.

  4. 4

    A healthcare provider requires a centralized data repository that can store massive volumes of unstructured medical imaging files while simultaneously supporting highly structured, ACID-compliant SQL transactions for patient billing records. Which of the following data repository architectures is BEST suited for this requirement?

    Show answer details

    Correct answer: C

    A Data Lakehouse is the correct answer. It is a modern architecture that combines the flexibility and cost-efficiency of a data lake (ideal for unstructured data like medical images) with the ACID transactional capabilities and data management features of a traditional data warehouse (required for billing records). Data Lakes lack native ACID transactional support, and Data Warehouses are poorly suited for unstructured data storage. A Data Mart is just a subset of a data warehouse focused on a specific business line.

  5. 5

    While configuring a new database table to store digitally signed PDF contracts that average 50MB in size, a database administrator asks you to define the data type for the contract column. Which data type is the MOST appropriate for this field?

    Show answer details

    Correct answer: C

    BLOB (Binary Large Object) is the correct data type for storing binary data such as images, multimedia, or PDF files. CLOB (Character Large Object) is used for massive amounts of text data (like large XML or JSON documents), not binary files. String is used for standard, short alphanumeric text. GUID is used for generating unique identifiers, not for file storage.

  6. 6

    An e-commerce company is modernizing its analytics capabilities. The project sponsor wants to implement technologies that can automatically summarize thousands of long-form customer reviews into single-paragraph insights, and also create realistic, synthetic demographic data to safely test a new application without exposing real PII. Which TWO of the following AI concepts should the team implement? (Select TWO)

    Show answer details

    Correct answer: B, D

    Natural Language Processing (NLP) is required to understand, interpret, and summarize the human language in the customer reviews. Generative AI is required to create the realistic, synthetic demographic data for testing. RPA is used for automating repetitive, rule-based tasks (like data entry), not generating text or synthetic data. Predictive analytics forecasts future outcomes based on historical data. Descriptive analytics summarizes what happened in the past.

    Generative AI is required to create the realistic, synthetic demographic data for testing. Natural Language Processing (NLP) is required to understand, interpret, and summarize the human language in the customer reviews. RPA is used for automating repetitive, rule-based tasks (like data entry), not generating text or synthetic data. Predictive analytics forecasts future outcomes based on historical data. Descriptive analytics summarizes what happened in the past.

Create an account to continue.