PDE Practice Questions
Prepare for PDE with more than an answer.
Unlock the full exam and previous versions
- v1Google Cloud Professional Data Engineer 145 questions Locked
- PDELegacy Professional Data Engineer 386 questions Current
- 1
Your company is using WILDCARD tables to query data across multiple tables with similar names. The SQL statement is currently failing with the following error:
Syntax error: Expected end of statement but got at [4:11]
SELECT age
FROM
bigquery-public-data.noaa_gsod.gsod
WHERE
age != 99 AND_TABLE_SUFFIX = ‘1929’
ORDER BY
age DESCWhich table name will make the SQL statement work correctly?
Show answer details
Correct answer: D
D-Reference: httosV/cloud.google.com/bigquery/docs/wildcard-tables
- 2
You work on a regression problem in a natural language processing domain, and you have 100M labeled exmaples in your dataset. You have randomly shuffled your data and split your dataset into train and test samples (in a 90/10 ratio). After you trained the neural network and evaluated your model on a test set, you discover that the root-mean-squared error (RMSE) of your model is twice as high on the train set as on the test set. How should you improve the performance of your model?
Show answer details
Correct answer: C
C
- 3
You need to create a new transaction table in Cloud Spanner that stores product sales data. You are deciding what to use as a primary key. From a performance perspective, which strategy should you choose?
Show answer details
Correct answer: C
Reference: https://www.uuidgenerator.net/version4
- 4
You need to copy millions of sensitive patient records from a relational database to BigQuery. The total size of the database is 10 TB. You need to design a solution that is secure and time-efficient.
What should you do?Show answer details
Correct answer: A
A
- 5
An e-commerce platform uses Cloud Composer to orchestrate a complex daily data processing workflow. The workflow involves several Dataproc jobs, BigQuery queries, and data validation checks. The platform team wants to implement a CI/CD process to automatically deploy changes to their Airflow DAGs from a GitHub repository to the Composer environment. Which of the following approaches is the recommended best practice for achieving this?
graph TD A[Developer pushes to main branch] --> B{GitHub Actions}; B --> C[Sync DAGs to Composer GCS Bucket]; C --> D[Cloud Composer picks up new DAG];Show answer details
Correct answer: B
The standard and recommended best practice for managing Cloud Composer DAGs is to treat them as code artifacts. A CI/CD pipeline (using Cloud Build, GitHub Actions, Jenkins, etc.) should be set up to automatically test and sync the contents of the repository's DAGs folder with the
/dagsfolder in the Cloud Composer environment's associated Cloud Storage bucket. Composer automatically detects and loads changes from this bucket. This approach provides version control, automated testing, and reliable deployments. - 6
You need to migrate a 2TB relational database to Google Cloud Platform. You do not have the resources to significantly refactor the application that uses this database and cost to operate is of primary concern.
Which service do you select for storing and serving your data?Show answer details
Correct answer: D
D
- 7
You are migrating your data warehouse to BigQuery. You have migrated all of your data into tables in a dataset. Multiple users from your organization will be using the data. They should only see certain tables based on their team membership. How should you set user permissions?
Show answer details
Correct answer: A
A
- 8
You are managing a Cloud Dataproc cluster. You need to make a job run faster while minimizing costs, without losing work in progress on your clusters. What should you do?
Show answer details
Correct answer: B
Reference: cloud.google.com/dataproc/docs/con">https://cloud.google.com/dataproc/docs/concepts/configuring-clusters/flex
- 9
You have a data stored in BigQuery. The data in the BigQuery dataset must be highly available. You need to define a storage, backup, and recovery strategy of this data that minimizes cost. How should you configure the BigQuery table?
Show answer details
Correct answer: C
C
- 10
You work for an economic consulting firm that helps companies identify economic trends as they happen. As part of your analysis, you use Google BigQuery to correlate customer data with the average prices of the 100 most common goods sold, including bread, gasoline, milk, and others. The average prices of these goods are updated every 30 minutes. You want to make sure this data stays up to date so you can combine it with other data in BigQuery as cheaply as possible. What should you do?
Show answer details
Correct answer: B
B
