Skip to content

Google Cloud Professional Data Engineer Practice Questions

Prepare for GCP-PDE with more than an answer.

145 questions in the full set12 sample questionsUpdated Mar 12, 2026

Unlock the full exam and previous versions

  • v1Google Cloud Professional Data Engineer 145 questions Current
  • PDELegacy Professional Data Engineer 386 questions Locked
Exam fee
$200 USD
Time limit
120 minutes
Questions on the exam
40-50
Passing score
~70%
Level
Professional
Valid for
2 years
Domains covered on the exam 5
  1. Designing data processing systems22%
  2. Ingesting and processing the data25%
  3. Storing the data20%
  4. Preparing and using data for analysis15%
  5. Maintaining and automating data workloads18%
  1. 1

    Case Study Scenario:

    A retail giant 'ShopSmart' is migrating its on-premises Oracle Exadata warehouse to Google Cloud. They have 2 PB of historical data and generate 5 TB of new transactional data daily.

    Requirements:

    1. The migration must be completed within 2 months.
    2. The cloud data warehouse must support standard SQL and scale to petabytes.
    3. Business Intelligence users currently use Looker and need sub-second response times for dashboards.
    4. Data sovereignty laws require that customer data from Europe stays in the europe-west3 region.

    Which storage and migration strategy meets these needs?

    Show answer details

    Correct answer: A

    BigQuery meets the SQL and scale requirements. Transfer Appliance is ideal for moving 2 PB quickly when bandwidth is a bottleneck (offline transfer). Datastream provides CDC for the 5 TB daily change. Regional datasets satisfy sovereignty. BI Engine specifically addresses the sub-second Looker performance requirement.

  2. 2

    You are designing a solution to replicate data from an on-premises MySQL database to BigQuery for real-time analytics. The solution must capture data changes (Inserts, Updates, Deletes) with minimal latency and minimal impact on the source database performance.

    Which Google Cloud service combination is best suited for this task?

    Show answer details

    Correct answer: A

    Datastream is a serverless Change Data Capture (CDC) and replication service. It reads directly from MySQL binary logs (minimizing impact compared to polling queries) and has a native integration to replicate data directly into BigQuery.

  3. 3

    Your startup app uses Cloud Firestore for its mobile backend. You need to perform complex analytical queries on this data, joining it with marketing data stored in CSV files on Cloud Storage. You want to minimize the engineering effort required to move or transform data.

    What is the most efficient solution?

    Show answer details

    Correct answer: A

    BigQuery supports querying Cloud Storage directly using external tables. For Firestore, you can either use the export-to-BigQuery extension or (more recently) BigLake/Object tables concepts. However, the classic pattern for 'minimal effort' is loading/federating the CSVs and using the Firestore-to-BigQuery extension to keep a copy synced, or exporting Firestore to GCS and querying both as external tables. The most direct 'query without moving' for Firestore is limited, but the extension is the standard answer for 'analyzing Firestore in BQ'.

  4. 4

    You are designing a data ingestion architecture for a global logistics company. The system must ingest telemetry data from millions of IoT devices installed in delivery trucks. The data arrives in bursts, and the order of events is critical for calculating accurate delivery estimates. However, due to intermittent connectivity, some devices may send data hours late. You need a solution that ensures strictly ordered processing per truck ID while handling high throughput and potential duplicates.

    Which architecture should you implement?

    Show answer details

    Correct answer: A

    Pub/Sub ordering keys ensure that messages with the same key (truck ID) are delivered to subscribers in the order they were published, which is critical for this scenario. Dataflow's streaming engine with 'exactly-once' processing handles the deduplication and ensures reliable processing of the ordered stream. Standard Pub/Sub without ordering keys does not guarantee order, and Kafka on Compute Engine adds unnecessary operational overhead.

  5. 5

    A financial institution requires a new data warehouse solution. The security team mandates that all data stored must be encrypted using keys that the institution manages and rotates annually. Additionally, specific columns containing PII (Personally Identifiable Information) like Social Security Numbers must be restricted so that only the 'HR-Auditors' group can view the plaintext values, while data analysts see masked data. You plan to use BigQuery.

    Which combination of features should you configure?

    Show answer details

    Correct answer: A, B

    To meet the requirement of managing and rotating keys, Customer-Managed Encryption Keys (CMEK) via Cloud KMS is the correct approach. For column-level visibility control, BigQuery Policy Tags (Data Catalog) with Dynamic Data Masking is the standard solution. Authorized Views could work for row-level or table-level access but are less efficient for granular column masking compared to Policy Tags. Select TWO.

    Policy tags allow for fine-grained access control and dynamic masking based on user roles.

  6. 6

    Your team is migrating a legacy Hadoop workload to Google Cloud. The workload currently runs on an on-premises cluster and processes 50 TB of log data daily using Spark jobs. The processing demand is highly variable: it spikes significantly between 2 AM and 6 AM and is idle for long periods. You need to minimize costs while ensuring the jobs complete within the 4-hour window.

    What architecture should you design?

    Show answer details

    Correct answer: A

    For variable workloads that are idle for long periods, ephemeral Dataproc clusters are the most cost-effective solution. You only pay for the compute resources while the job is running. Persistent clusters would incur costs during the idle 20 hours. Migrating to Dataflow would require rewriting the Spark code, which is not specified as a desire. Compute Engine requires manual management and is less efficient.

Create an account to continue.