PCDOE Professional Cloud DevOps Engineer Practice Questions
Prepare for PCDOE with more than an answer.
- Exam fee
- $200 USD
- Level
- Professional
- Valid for
- 2 years
Domains covered on the exam 5
- Bootstrapping and maintaining a Google Cloud organization17%
- Building and implementing CI/CD pipelines for applications and infrastructure27%
- Applying site reliability engineering practices to applications23%
- Implementing observability practices20%
- Optimizing performance and troubleshooting13%
- 1
You are deploying a Cloud Run service that needs to connect to a Cloud SQL database using a database password. To follow security best practices, you must avoid storing the password in the container image, in build logs, or as a plaintext environment variable. What is the recommended method to securely provide the password to the Cloud Run service at runtime?
Show answer details
Correct answer: C
This is the most secure and manageable solution. Secret Manager is designed for storing secrets. Cloud Run has a direct integration with Secret Manager that allows you to mount a secret's value as a file in an in-memory volume or expose it directly as an environment variable. This process is handled by the Google Cloud infrastructure at runtime, ensuring the secret is never exposed in the container image, build process, or service configuration manifests. The Cloud Run service's runtime service account is granted IAM permission to access only the specific secret it needs.
- 2
Your team's primary service has a 99.9% availability SLO. Due to a series of minor incidents, the service has completely consumed its error budget for the month with two weeks remaining. The product team wants to launch a new high-risk, high-reward feature immediately. According to SRE principles, what is the most appropriate course of action?
Show answer details
Correct answer: C
The error budget is a data-driven tool for balancing reliability and feature velocity. When the budget is exhausted, the pre-agreed policy should be that the risk of further instability is too high. The team's focus must pivot from releasing new features (which introduce risk) to reliability-enhancing work. This could include fixing bugs, improving monitoring, automating manual tasks, or addressing root causes from recent incidents. The feature launch should be postponed until an adequate error budget is available again.
- 3
The primary command-line interface (CLI) tool for interacting with and managing resources on Google Cloud Platform is known as the ______ CLI.
Show answer details
Correct answer: C
The
gcloudcommand-line tool is the main CLI for Google Cloud. It is part of the Google Cloud SDK and can be used to manage a wide range of products and services, including Compute Engine, GKE, Cloud Run, and IAM. - 4
An enterprise platform team manages dozens of GKE clusters for various development teams. They need to enforce a set of consistent security policies across all clusters, such as requiring all container images to be from a trusted Artifact Registry repository and disallowing pods from running with root privileges. What is the recommended Google Cloud solution to manage and enforce these policies at scale?
Show answer details
Correct answer: B
GKE Enterprise (formerly Anthos) includes Policy Controller, which is built on the open-source Open Policy Agent (OPA) Gatekeeper project. It allows you to define policies as custom resources (constraints) and apply them centrally to a 'fleet' of GKE clusters. This provides a scalable, declarative, and GitOps-friendly way to enforce security and configuration consistency across an entire enterprise without manual intervention on individual clusters.
- 5
During a postmortem for a cascading failure that impacted multiple services, your team identifies a lack of proper isolation as a key contributing factor. The failure of a non-critical downstream service caused a resource bottleneck that brought down a critical upstream service. Which THREE of the following SRE practices are most effective at preventing or mitigating such cascading failures? (Select THREE)
flowchart TD A[Start Incident] --> B{Is critical service A affected?} B -- Yes --> C[Page Primary On-Call] B -- No --> D{Is non-critical service B affected?} C --> E[Execute Playbook A] D -- Yes --> F[Page Secondary On-Call] F --> G[Execute Playbook B] E --> H{Apply Mitigation} G --> H H --> I[Resolve Incident] I --> J[Conduct Postmortem]Show answer details
Correct answer: A, C, E
A circuit breaker wraps network calls. If a downstream service starts failing or timing out, the breaker 'trips' and immediately fails subsequent calls without waiting for a timeout. This prevents the upstream service from consuming resources (like threads or connections) waiting for a failing dependency, thus protecting it from the downstream failure.
The bulkhead pattern partitions a service's resources (e.g., connection pools, thread pools) so that a failure in one area doesn't exhaust all resources and bring down the entire service. For example, calls to service A use one connection pool, and calls to service B use another. If service B fails, it only exhausts its own pool, leaving the pool for service A intact.
Load shedding is the practice of intentionally dropping excess or low-priority requests when a system is overloaded. This ensures that the system can continue to serve high-priority traffic and remain stable, rather than failing completely for all users. It's a key strategy for preventing total collapse during unexpected load spikes or resource contention.
- 6
A financial services company is building a CI/CD pipeline using Cloud Build and Cloud Deploy. To comply with internal security policies, they must ensure that only container images that have passed all vulnerability scans and integration tests are deployable to their production GKE clusters. Furthermore, the mechanism enforcing this must be resistant to tampering, even by users with project owner roles. Which combination of services should be implemented to meet these requirements?
Show answer details
Correct answer: C
This is the most secure and tamper-resistant solution. Binary Authorization is a service specifically designed to enforce deployment-time policies on GKE. By creating a cryptographic attestation (a signature proving the image passed all checks) in a secure build environment and requiring that attestation in the GKE cluster's policy, you create a strong, auditable link. This enforcement happens at the GKE control plane level and cannot be easily bypassed, even by project owners, without modifying the Binary Authorization policy itself, which can be tightly controlled.
- 7
You are the SRE lead for a critical order processing service composed of three microservices: an API Gateway, a Processing Service, and a Database Writer. The overall service is considered successful only if an order is accepted by the gateway, fully processed by the service, and successfully written to the database. The individual SLOs are: Gateway (99.95% availability), Processing Service (99.9% availability), and Database Writer (99.99% availability). Assuming these components fail independently, what is the correct composite SLO for the end-to-end user journey?
Show answer details
Correct answer: C
For a user journey that depends on multiple components in series, the overall availability is the product of the individual component availabilities. Each component represents a potential point of failure for the entire transaction. Therefore, the composite SLO is calculated by multiplying the individual SLOs: 0.9995 * 0.999 * 0.9999 ≈ 0.9984, or 99.84%. This demonstrates that serial dependencies decrease overall reliability.
- 8
You are managing a multi-tenant GKE cluster where each tenant application runs in its own namespace. One tenant deployed a custom Fluentd configuration as a DaemonSet to forward their specific application logs to an external analytics service. Shortly after, your central Cloud Logging view stops receiving any logs from the nodes where this tenant's pods are running. Logs from other applications on the same nodes are also missing. What is the most likely cause of this issue?
Show answer details
Correct answer: B
GKE uses a managed Fluentd DaemonSet as its default logging agent on each node. When a user deploys another log-forwarding DaemonSet, especially one that also tries to read from standard log locations like
/var/log, it can create resource contention or configuration conflicts. The most common issue is that the custom agent 'steals' the log files or interferes with the default agent's ability to read them, effectively stopping the flow of logs to Cloud Logging for that entire node. - 9
A development team is starting a new project on Google Cloud and needs to manage their infrastructure as code. The team has extensive experience with Kubernetes YAML but is new to cloud infrastructure provisioning. They want a tool that allows them to define their GCP resources using a Kubernetes-style, declarative syntax and manage them through their GKE cluster. Which IaC tool is the best fit for this team's requirements?
Show answer details
Correct answer: C
Config Connector is specifically designed for this use case. It is a Kubernetes addon that allows you to manage GCP resources (like Pub/Sub topics, Cloud SQL instances, etc.) as Kubernetes Custom Resource Definitions (CRDs). This enables teams familiar with
kubectland YAML manifests to use the same tools and workflows to manage both their applications and their underlying cloud infrastructure, fitting their existing skillset perfectly. - 10
An e-commerce application running on GKE is experiencing intermittent latency spikes during checkout. The system consists of a frontend service, a cart service, and an inventory service. To begin troubleshooting, you need to understand the flow of requests and identify which service call is introducing the most latency. What are the TWO most effective initial steps to take? (Select TWO)
sequenceDiagram participant User participant Frontend participant CartSvc as Cart Service participant InvSvc as Inventory Service User->>Frontend: POST /checkout Frontend->>CartSvc: GET /cart/items CartSvc-->>Frontend: Cart Items Frontend->>InvSvc: POST /inventory/reserve InvSvc-->>Frontend: Reservation Confirmed Frontend-->>User: Checkout CompleteShow answer details
Correct answer: A, D
While Cloud Trace is excellent for inter-service latency, Cloud Profiler is crucial for intra-service latency. If a trace shows a specific service is slow, Profiler is the next step to pinpoint the exact function or code path within that service causing high CPU usage or lock contention, which are common causes of latency.
Cloud Trace is the primary tool for understanding latency in a distributed microservices architecture. By analyzing the trace waterfall for a slow request, you can visually identify which service-to-service call is taking the longest time, immediately narrowing down the scope of the investigation.
