DP-100 Practice Questions
Prepare for DP-100 with more than an answer.
- Exam fee
- $165 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 4
- Design and prepare a machine learning solution22.5%
- Explore data, and run experiments22.5%
- Train and deploy models27.5%
- Optimize language models for AI applications27.5%
- 1
A data scientist is using Automated ML for a time-series forecasting task to predict weekly sales. The dataset contains several years of historical data. To improve model accuracy, they want to incorporate features based on the week of the year and the day of the week. How can they ensure AutoML automatically generates these time-based features?
Show answer details
Correct answer: B
For forecasting tasks, AutoML has powerful built-in time-series featurization capabilities. By correctly identifying the time column in the forecasting settings and leaving featurization on 'auto' (the default), AutoML will automatically generate a rich set of features from the date/time column, such as year, month, day, day of week, week of year, and more.
- 2
You are building a RAG solution and have configured an Azure AI Search index as your vector store. You want to improve the relevance of search results by having the system understand the user's intent rather than just matching keywords. Which Azure AI Search feature should you enable and configure?
Show answer details
Correct answer: C
The semantic ranker (also known as semantic search) is a premium feature in Azure AI Search that uses deep learning models to understand the contextual meaning and intent behind a query. It re-ranks the initial set of results from keyword or vector search based on semantic relevance, significantly improving the quality of results for RAG applications.
- 3
A batch endpoint has been deployed to score customer data daily. The scoring job is failing, and the logs for the job run show an
OutOfMemoryError. The job runs on a compute cluster withStandard_DS4_v2nodes. What is the most likely cause and the most appropriate solution to fix this issue?Show answer details
Correct answer: B
An
OutOfMemoryErrorin a batch scoring job typically occurs when the amount of data being processed in a single batch exceeds the available RAM on the worker node. Themini_batch_sizeparameter controls how many files or records are processed at once. By reducing this value, less data is loaded into memory simultaneously, which should resolve the memory issue without needing to change the compute hardware. - 4
A data science team uses MLflow extensively to track experiments. They want to programmatically identify the best run from a specific experiment based on the lowest 'Root Mean Squared Error' and then register the corresponding model. Which tools are required to accomplish this? (Select TWO)
Show answer details
Correct answer: A, B
The
mlflow.search_runsfunction is used to query past runs. It allows you to filter by experiment ID and order the results by a specific metric (e.g., 'metrics.rmse' in ascending order) to find the best performing run.Once the best run is identified, the
mlflow.register_modelfunction is used to take the model artifact from that run (specified by its run URI) and register it in the MLflow Model Registry with a given name. - 5
A retail company wants to fine-tune a foundation model to act as a customer support chatbot that uses a specific, friendly brand voice. They have a dataset of 10,000 high-quality conversation logs between expert support agents and customers. What is the most appropriate optimization approach?
Show answer details
Correct answer: C
Fine-tuning is the ideal approach when the goal is to teach a model a new skill, style, or format that is consistent across many interactions. By training the model on a large, high-quality dataset of expert conversations, you directly adjust the model's weights to adopt the specific brand voice and response patterns demonstrated in the logs. RAG is for providing knowledge, and prompt engineering is less effective for instilling a consistent, nuanced style.
- 6
You need to create a data asset in your Azure Machine Learning workspace that points to a specific folder in an Azure Data Lake Storage Gen2 account. The data in this folder will be updated periodically, and you want your training scripts to always use the latest version of the data without needing to change the script. How should you create and reference the data asset?
Show answer details
Correct answer: C
Azure ML data assets support versioning. You can create new versions as the underlying data changes. By referencing the asset in your script using the
@latesttag (e.g.,azureml:my_data@latest), you ensure that the job always resolves to the most recently created version of that asset, achieving the goal without code modifications. - 7
Case Study:
A travel company, GoExplore, is building a sophisticated customer support agent using Azure AI. The agent needs to handle complex, multi-turn conversations. A typical user interaction involves identifying the customer, retrieving their booking details from a CRM API, checking flight status from an external airline API, and then providing a summary and potential rebooking options.
The current implementation is struggling. The agent sometimes fails to call the APIs in the correct order, and debugging the conversational flow is difficult. The team needs a solution that provides robust orchestration, clear visualization of the logic flow, and better management of API connections and prompts.
Which Azure Machine Learning technology should the team adopt to structure, orchestrate, and debug their support agent?
flowchart TD A(User Query) --> B{Intent Classification}; B -->|'Get Booking'| C[Call CRM API]; B -->|'Check Status'| D[Call Airline API]; C --> E{Booking Found?}; E -->|Yes| F[Extract Flight Info]; F --> D; D --> G[Generate Summary]; G --> H(Agent Response); E -->|No| I[Ask for Booking ID]; I --> A;Show answer details
Correct answer: D
Prompt flow is specifically designed for this type of complex LLM orchestration. Its visual graph allows developers to clearly define and visualize the flow of logic, including conditional branches. It has native support for Python tools to encapsulate API calls, LLM nodes for prompt execution, and a secure 'Connections' system to manage secrets like API keys. The built-in tracing and debugging capabilities would directly address the team's difficulties in understanding and troubleshooting the conversational flow.
- 8
A financial services company is developing a Retrieval-Augmented Generation (RAG) solution to answer questions about internal compliance documents. The documents are a mix of short policy statements (1-2 paragraphs) and long procedural guides (10-20 pages). The goal is to ensure that answers are precise and source attribution is accurate. Which data preparation strategy is most suitable for this scenario?
Show answer details
Correct answer: B
A recursive character text splitter is ideal for documents with varied structure. It attempts to split along semantic boundaries (paragraphs, sections) first before falling back to smaller units. This preserves the context within chunks, which is crucial for accurate retrieval from both short policies and long guides. A moderate chunk size with overlap ensures that sentences or ideas are not awkwardly split between chunks.
- 9
A data science team is using Azure Machine Learning pipelines to orchestrate a complex training workflow. A custom component in the middle of the pipeline frequently fails due to transient network issues when accessing an external data source. The team wants to make the pipeline more resilient without modifying the component's internal code. How should they configure the pipeline job to handle these intermittent failures?
Show answer details
Correct answer: C
Azure Machine Learning components can be configured with retry settings directly in their YAML definition. By adding a
retry_settingsblock with properties likecount,delay, andbackoff, you instruct the pipeline orchestrator to automatically re-run the component if it fails. This is the correct approach for handling transient errors without altering the component's source code. - 10
You are designing a secure environment for a multi-team data science project. You need to ensure that each team can manage its own compute resources and data assets but cannot access the resources of other teams. All teams must use a centrally-managed set of curated Docker environments and foundation models. Which combination of Azure Machine Learning features should you use? (Select TWO)
Show answer details
Correct answer: B, C
Creating separate workspaces provides the strongest isolation boundary. Each team can manage its own compute, data, and experiments independently, satisfying the requirement that they cannot access each other's resources.
An Azure Machine Learning registry is designed for sharing assets like models, environments, and components across multiple workspaces. This allows a central MLOps team to manage and distribute curated assets to all the individual team workspaces.
