DP-420 Practice Questions
Prepare for DP-420 with more than an answer.
- Exam fee
- $165 USD
- Level
- Specialty
- Valid for
- 2 years
Domains covered on the exam 5
- Design and implement data models40%
- Design and implement data distribution7%
- Integrate an Azure Cosmos DB solution8%
- Optimize an Azure Cosmos DB solution17%
- Maintain an Azure Cosmos DB solution28%
- 1
You are designing a cost-effective solution for storing user audit logs in Azure Cosmos DB for NoSQL. The logs must be retained for 90 days for compliance reasons and then automatically deleted. The volume of logs is highly unpredictable, ranging from very low to extremely high throughout the day. Queries on the logs are infrequent. Which combination of Azure Cosmos DB capacity mode and feature should you use to meet these requirements most cost-effectively?
Show answer details
Correct answer: C
Serverless capacity mode is ideal for unpredictable, spiky workloads with infrequent traffic, as you only pay for the RUs consumed and storage used, without any minimum charge for provisioned throughput. Time-to-Live (TTL) is the native Cosmos DB feature for automatically deleting documents after a specified duration. Setting the default TTL on the container to 7,776,000 seconds (90 days) provides a free and efficient mechanism for data deletion, perfectly matching the requirements.
- 2
You have an Azure Cosmos DB for NoSQL account with a container that is configured for multi-region writes. The conflict resolution policy is set to Last Writer Wins (LWW), using the default
_tstimestamp property. An application instance in 'East US' and another in 'West Europe' simultaneously update the same document. The write from 'East US' has a timestamp of1662402705, and the write from 'West Europe' has a timestamp of1662402703. Which update will be persisted as the final version of the document?Show answer details
Correct answer: C
The Last Writer Wins (LWW) conflict resolution policy resolves conflicts based on a specified timestamp path. When using the default
_tsproperty, which is a server-side timestamp, the update with the highest (i.e., most recent) timestamp value will win the conflict. In this case,1662402705(East US) is greater than1662402703(West Europe), so the update from East US will be persisted. - 3
A developer is querying a container that stores product information. Each document has a
tagsproperty which is an array of strings, for example:"tags": ["electronics", "wearable", "smartwatch"]. The developer needs to write a query that returns all documents where thetagsarray contains the string "wearable". Which built-in function is the most efficient for this purpose?Show answer details
Correct answer: B
The
ARRAY_CONTAINS()function is specifically designed and optimized for checking if an array contains a specific value. It is the correct and most efficient function to use for this type of query. For example:SELECT * FROM c WHERE ARRAY_CONTAINS(c.tags, "wearable"). - 4
Case Study: Contoso Fitness Wearables
Company Background:
Contoso Fitness Wearables is a rapidly growing company that produces smart fitness devices. They are launching a new global cloud-native application to track user activities, store biometric data, and provide personalized health insights. The application is expected to serve millions of users across North America, Europe, and Southeast Asia. Low latency for both reads and writes is a critical success factor to ensure a responsive user experience.Current Situation:
The development team has chosen Azure Cosmos DB for NoSQL as the primary database. They have designed a data model where each user has a document containing their profile and an embedded array of recent activities. The proposed partition key is/userId. The application will be deployed to Azure regions in Central US, West Europe, and Southeast Asia.Requirements:
- Global Low Latency: Users in all three regions must experience fast read and write performance, ideally under 10ms.
- High Availability: The application must be resilient to regional outages. An automatic failover mechanism is required.
- Data Consistency: User profile updates (e.g., changing subscription tier) must be strongly consistent. However, activity tracking data can tolerate a few seconds of replication lag to prioritize write availability.
- Operational Efficiency: The solution for handling different consistency requirements must be implemented within the application code using the SDK, without requiring separate containers for profile and activity data.
Problem:
You are the lead architect responsible for finalizing the Azure Cosmos DB configuration. You need to devise a strategy that meets all the specified requirements for latency, availability, and mixed consistency. Which configuration should you recommend?Show answer details
Correct answer: C
This solution correctly addresses all requirements. 1) Enabling multi-region writes provides low-latency writes in all three regions. 2) Configuring multiple regions with an automatic failover policy meets the high availability need. 3) Setting a relaxed default consistency like Bounded Staleness or Session is good for general use cases like activity tracking. 4) Crucially, the SDK allows developers to override the consistency level on a per-request basis. This allows the application to elevate the consistency to Strong specifically for critical user profile updates, while using the more relaxed (and more available) default for other operations, all within a single container as required.
- 5
You are analyzing query performance in your Azure Cosmos DB for NoSQL solution. You execute a query and retrieve the diagnostics, which include the query metrics. The metrics show a 'Retrieved Document Count' of 5,000 but an 'Output Document Count' of 10. What does this discrepancy indicate about the query's efficiency?
flowchart LR A[Query Engine] --> B{Index Seek/Scan}; B --> C[Load 5,000 Documents]; C --> D{Apply Filter in Backend}; D --> E[Output 10 Documents]; subgraph Legend direction TB F(Retrieved Document Count = 5000) G(Output Document Count = 10) endShow answer details
Correct answer: B
A large gap between 'Retrieved Document Count' and 'Output Document Count' is a key indicator of an inefficient query. It means the index was able to narrow down the search to 5,000 documents, but the query engine then had to load all 5,000 into memory to apply additional filter conditions that were not served by the index, ultimately discarding 4,990 of them. This consumes significantly more RUs than necessary. The solution is often to create a more specific or composite index that better matches the query's filter criteria.
- 6
A financial services company is using Azure Cosmos DB for NoSQL to store transaction data. The container is partitioned by
/transactionId. During a performance audit, you observe that queries filtering on/transactionDateare consuming a high number of RUs and are identified as cross-partition queries. The development team wants to add a composite index to improve performance for queries that filter by both/transactionDateand/transactionType. The current indexing policy is the default. What is the most likely outcome of adding a composite index for (/transactionDate,/transactionType)?Show answer details
Correct answer: B
A composite index can optimize the filtering and sorting of cross-partition queries, thereby reducing their RU cost. However, it does not change the fundamental nature of the query. Since the partition key (
/transactionId) is not included in the query's WHERE clause, the query must still be fanned out across all partitions to find the relevant documents. The composite index makes the process of finding matching documents within each partition much more efficient. - 7
You are designing a multi-tenant SaaS application on Azure Cosmos DB for NoSQL. Each tenant's data must be logically isolated. The application has a high volume of small, frequent writes and reads. To optimize costs, you provision throughput at the database level and use a shared container for all tenants. You choose
/tenantIdas the partition key. A new requirement mandates that within each tenant's data, all user profiles must have a uniqueemailAddress. How should you enforce this new uniqueness constraint?Show answer details
Correct answer: D
A unique key constraint in Azure Cosmos DB is always scoped to a logical partition. By defining a unique key policy on
/emailAddressand having/tenantIdas the partition key, you enforce thatemailAddressmust be unique for all documents that share the sametenantId. This perfectly matches the requirement for tenant-level data isolation. - 8
A global e-commerce platform uses Azure Cosmos DB for NoSQL with multi-region writes enabled to reduce latency for a worldwide user base. The consistency level is set to Session. To handle concurrent updates to a user's shopping cart, a custom conflict resolution policy has been implemented using a merge stored procedure. The policy is designed to merge the items from conflicting writes. During a sales event, users in Europe report that items they add to their cart are occasionally disappearing. Users in North America do not report this issue. The primary write region is North America. What is the most probable cause of this issue?
Show answer details
Correct answer: B
When using a custom conflict resolution policy with a merge stored procedure, the stored procedure must be created in each region where writes can occur. If the stored procedure is only registered in the North America region, conflicts originating from writes in Europe will fail to be resolved correctly, leading to data loss. The conflict resolution mechanism will fall back to a default behavior which might not align with the intended merge logic.
- 9
You are building an event-sourcing solution where all changes to application state are captured as a sequence of events. You use Azure Cosmos DB for NoSQL and an Azure Function with a Cosmos DB trigger to process these events. The function archives events to cold storage and updates several materialized views in other containers. During testing, you notice that if the Azure Function fails while processing a batch of changes, it re-processes the same batch again upon restart, leading to duplicate data in the materialized views. Which two actions should you take to make the function idempotent and prevent data duplication? (Select TWO)
Show answer details
Correct answer: A, D
The
_lsnproperty is a unique, monotonically increasing number for each change. By tracking the last successfully processed_lsnfor each partition in your destination (or a separate state store), your function can check if an incoming event's_lsnhas already been processed, thus achieving idempotency.Making the write operations to the materialized views idempotent is crucial. Using an upsert (
CreateItemAsyncorReplaceItemAsynclogic) ensures that if an event is processed twice, the second operation simply overwrites the existing data with the same information, rather than creating a duplicate entry. - 10
You are managing a large Azure Cosmos DB for NoSQL container that stores user session data. The container is configured with the default indexing policy (automatic indexing of all properties). You identify that a specific property,
sessionTrace, contains large JSON blobs used only for debugging and is never queried. To reduce storage costs and improve write performance, you decide to exclude this property from indexing. Which JSON snippet represents the correct indexing policy to achieve this?Show answer details
Correct answer: B
This policy correctly uses the default behavior of indexing everything (
includedPathswith/*) and then specifically carves out an exception. TheexcludedPathsarray with the path/sessionTrace/*tells Cosmos DB to not index thesessionTraceproperty or any of its sub-properties. The*is a wildcard that matches any element beyond thesessionTraceproperty.
