DP-100 Practice Questions
Prepare for DP-100 with more than an answer.
- Exam fee
- $165 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 4
- Design and prepare a machine learning solution22.5%
- Explore data, and run experiments22.5%
- Train and deploy models27.5%
- Optimize language models for AI applications27.5%
- 1
A batch endpoint has been deployed to score customer data daily. The scoring job is failing, and the logs for the job run show an
OutOfMemoryError. The job runs on a compute cluster withStandard_DS4_v2nodes. What is the most likely cause and the most appropriate solution to fix this issue?Show answer details
Correct answer: B
An
OutOfMemoryErrorin a batch scoring job typically occurs when the amount of data being processed in a single batch exceeds the available RAM on the worker node. Themini_batch_sizeparameter controls how many files or records are processed at once. By reducing this value, less data is loaded into memory simultaneously, which should resolve the memory issue without needing to change the compute hardware. - 2
A data science team uses MLflow extensively to track experiments. They want to programmatically identify the best run from a specific experiment based on the lowest 'Root Mean Squared Error' and then register the corresponding model. Which tools are required to accomplish this? (Select TWO)
Show answer details
Correct answer: A, B
The
mlflow.search_runsfunction is used to query past runs. It allows you to filter by experiment ID and order the results by a specific metric (e.g., 'metrics.rmse' in ascending order) to find the best performing run.Once the best run is identified, the
mlflow.register_modelfunction is used to take the model artifact from that run (specified by its run URI) and register it in the MLflow Model Registry with a given name. - 3
A retail company wants to fine-tune a foundation model to act as a customer support chatbot that uses a specific, friendly brand voice. They have a dataset of 10,000 high-quality conversation logs between expert support agents and customers. What is the most appropriate optimization approach?
Show answer details
Correct answer: C
Fine-tuning is the ideal approach when the goal is to teach a model a new skill, style, or format that is consistent across many interactions. By training the model on a large, high-quality dataset of expert conversations, you directly adjust the model's weights to adopt the specific brand voice and response patterns demonstrated in the logs. RAG is for providing knowledge, and prompt engineering is less effective for instilling a consistent, nuanced style.
- 4
You need to create a data asset in your Azure Machine Learning workspace that points to a specific folder in an Azure Data Lake Storage Gen2 account. The data in this folder will be updated periodically, and you want your training scripts to always use the latest version of the data without needing to change the script. How should you create and reference the data asset?
Show answer details
Correct answer: C
Azure ML data assets support versioning. You can create new versions as the underlying data changes. By referencing the asset in your script using the
@latesttag (e.g.,azureml:my_data@latest), you ensure that the job always resolves to the most recently created version of that asset, achieving the goal without code modifications. - 5
Case Study:
A travel company, GoExplore, is building a sophisticated customer support agent using Azure AI. The agent needs to handle complex, multi-turn conversations. A typical user interaction involves identifying the customer, retrieving their booking details from a CRM API, checking flight status from an external airline API, and then providing a summary and potential rebooking options.
The current implementation is struggling. The agent sometimes fails to call the APIs in the correct order, and debugging the conversational flow is difficult. The team needs a solution that provides robust orchestration, clear visualization of the logic flow, and better management of API connections and prompts.
Which Azure Machine Learning technology should the team adopt to structure, orchestrate, and debug their support agent?
flowchart TD A(User Query) --> B{Intent Classification}; B -->|'Get Booking'| C[Call CRM API]; B -->|'Check Status'| D[Call Airline API]; C --> E{Booking Found?}; E -->|Yes| F[Extract Flight Info]; F --> D; D --> G[Generate Summary]; G --> H(Agent Response); E -->|No| I[Ask for Booking ID]; I --> A;Show answer details
Correct answer: D
Prompt flow is specifically designed for this type of complex LLM orchestration. Its visual graph allows developers to clearly define and visualize the flow of logic, including conditional branches. It has native support for Python tools to encapsulate API calls, LLM nodes for prompt execution, and a secure 'Connections' system to manage secrets like API keys. The built-in tracing and debugging capabilities would directly address the team's difficulties in understanding and troubleshooting the conversational flow.
- 6
A financial services company is developing a Retrieval-Augmented Generation (RAG) solution to answer questions about internal compliance documents. The documents are a mix of short policy statements (1-2 paragraphs) and long procedural guides (10-20 pages). The goal is to ensure that answers are precise and source attribution is accurate. Which data preparation strategy is most suitable for this scenario?
Show answer details
Correct answer: B
A recursive character text splitter is ideal for documents with varied structure. It attempts to split along semantic boundaries (paragraphs, sections) first before falling back to smaller units. This preserves the context within chunks, which is crucial for accurate retrieval from both short policies and long guides. A moderate chunk size with overlap ensures that sentences or ideas are not awkwardly split between chunks.
- 7
A data science team is using Azure Machine Learning pipelines to orchestrate a complex training workflow. A custom component in the middle of the pipeline frequently fails due to transient network issues when accessing an external data source. The team wants to make the pipeline more resilient without modifying the component's internal code. How should they configure the pipeline job to handle these intermittent failures?
Show answer details
Correct answer: C
Azure Machine Learning components can be configured with retry settings directly in their YAML definition. By adding a
retry_settingsblock with properties likecount,delay, andbackoff, you instruct the pipeline orchestrator to automatically re-run the component if it fails. This is the correct approach for handling transient errors without altering the component's source code. - 8
You are designing a secure environment for a multi-team data science project. You need to ensure that each team can manage its own compute resources and data assets but cannot access the resources of other teams. All teams must use a centrally-managed set of curated Docker environments and foundation models. Which combination of Azure Machine Learning features should you use? (Select TWO)
Show answer details
Correct answer: B, C
Creating separate workspaces provides the strongest isolation boundary. Each team can manage its own compute, data, and experiments independently, satisfying the requirement that they cannot access each other's resources.
An Azure Machine Learning registry is designed for sharing assets like models, environments, and components across multiple workspaces. This allows a central MLOps team to manage and distribute curated assets to all the individual team workspaces.
- 9
A data scientist is using the Azure Machine Learning SDK v2 to submit a hyperparameter tuning job for a classification model. The goal is to maximize the 'AUC_weighted' metric. The search space is large, and the compute budget is limited. They need to configure the sweep job to efficiently find good parameters by terminating underperforming runs early. Which early termination policy is most appropriate for this goal?
Show answer details
Correct answer: A
The Bandit policy is an aggressive early termination policy that terminates runs whose primary metric falls outside a specified slack factor/amount compared to the best-performing run. This is highly effective for efficiently exploring a large search space with a limited budget by quickly discarding unpromising trials.
- 10
You are developing a prompt flow that orchestrates multiple calls to a language model to generate a marketing campaign proposal. You need to ensure that the output of an early step, which generates a target audience description, is correctly passed as input to a later step that writes ad copy. Which Prompt flow feature allows you to define this data dependency?
Show answer details
Correct answer: C
In Prompt flow, you use Jinja2 templating syntax to reference the outputs of previous nodes. For example, in the ad copy node's prompt, you would write
{{generate_audience.output}}to insert the output from thegenerate_audiencenode. This creates the explicit data dependency and chains the steps together.
