Google Cloud Professional Data Engineer Practice Exam
Practice questions for the Google Cloud Professional Data Engineer certification: designing secure and compliant data systems with IAM, policy tags, row access policies, CMEK, Cloud EKM, VPC Service Controls and Sensitive Data Protection; reliability, data quality with Dataform assertions and Knowledge Catalog (formerly Dataplex Universal Catalog), and BigQuery managed disaster recovery; migrations with BigQuery Data Transfer Service, Database Migration Service, Datastream and Transfer Appliance; ingestion and processing with Pub/Sub, Dataflow and Apache Beam windowing, triggers and late data, Managed Service for Apache Spark (formerly Dataproc), Managed Service for Apache Kafka, Cloud Data Fusion, Dataform and the BigQuery Storage Write API; storage choices across BigQuery partitioning and clustering, BigLake and Apache Iceberg tables, Bigtable, Spanner, Cloud SQL and Cloud Storage; analysis with materialized views, BI Engine, Looker, data masking, BigQuery ML, Feature Store on Gemini Enterprise Agent Platform (formerly Vertex AI Feature Store), embeddings and vector search, BigQuery sharing (formerly Analytics Hub) and data clean rooms; and operations with BigQuery editions, reservations and quotas, Managed Service for Apache Airflow (formerly Cloud Composer), Workflows, CI/CD, Cloud Monitoring, time travel, snapshots and failover. Every question includes a written explanation.
100 questions · 12 free preview
Studying more than one? All Google Cloud exams for $29 · every exam for $79
Free sample questions
- Sample · question 1 · Sliding windows for moving averages
A card-fraud pipeline at Ivybridge Payments must compute, for every card, the average transaction amount over the previous 10 minutes, and refresh that figure every minute as events stream in through Dataflow. Which windowing strategy fits?
- A.Fixed windows of one minute
- B.Session windows with a 10-minute gap
- C.Fixed windows of 10 minutes
- D.Sliding windows 10 minutes long that start every minutecorrect
Why: Sliding (hopping) windows overlap: each one covers 10 minutes, and a new one begins every minute, so each card gets a fresh 10-minute average once a minute. One-minute fixed windows only see one minute of data, and 10-minute fixed windows update only once every 10 minutes. Session windows depend on gaps in activity, not a rolling time span.
Open this question on its own page → - Sample · question 2 · Turbo replication on dual-region buckets
Brackenridge Broadcasting stores newly recorded programs in a Cloud Storage bucket and needs every new object copied to a second region within 15 minutes so that a regional outage loses at most 15 minutes of recordings. Which bucket configuration meets this target?
- A.A dual-region bucket with turbo replication enabledcorrect
- B.A multi-region bucket with default replication
- C.A regional bucket with Object Versioning
- D.A regional bucket copied nightly to another region by Storage Transfer Service
Why: Turbo replication is offered on dual-region buckets and targets replicating 100% of newly written objects to both regions within 15 minutes. Default replication runs asynchronously and does not have a 15-minute target. Object Versioning keeps older versions in the same region, and a nightly copy could lose up to a day of recordings.
Open this question on its own page → - Sample · question 3 · Assured Workloads for sovereign controls
Kerensa Health, a European provider, must make sure that new projects for its analytics platform keep data in EU regions, limit support access to EU personnel and enforce these rules automatically as teams create resources. What should the cloud team set up?
- A.Labels on each project marking it as EU-only
- B.A VPC Service Controls perimeter around the analytics projects
- C.An Assured Workloads folder with an EU data residency and support control package, and create the analytics projects in itcorrect
- D.Instructions for teams to always choose EU locations when creating datasets
Why: Assured Workloads applies a control package to a folder, enforcing allowed locations through organization policies and providing personnel and support controls for the regime chosen. Labels and written guidance enforce nothing. VPC Service Controls limit data movement across a perimeter but do not restrict resource locations or support personnel.
Open this question on its own page → - Sample · question 4 · Data Validation Tool after migration
After moving 400 tables from an on-premises PostgreSQL warehouse to BigQuery, Holloway Insurance's auditors want evidence that every table arrived complete and unchanged, including row counts, column sums and row-level comparisons. Which approach provides this with the least custom code?
- A.Compare the storage size of each table in INFORMATION_SCHEMA.TABLE_STORAGE with the source
- B.Accept the migration job's success status as proof of completeness
- C.Run Google's open-source Data Validation Tool to compare counts, aggregates and row hashes between source and targetcorrect
- D.Run Knowledge Catalog (formerly Dataplex Universal Catalog) data profile scans on the BigQuery tables
Why: The Data Validation Tool connects to both systems and runs column, row-count, aggregate and row-hash validations, producing a report of matches and differences. Storage size differs between engines because of compression and formats. A job's success status does not prove the data matches, and profiling the target alone has nothing to compare against.
Open this question on its own page → - Sample · question 5 · Pub/Sub Cloud Storage subscription archive
Compliance at Avonmore Logistics requires a raw copy of every message from a Pub/Sub topic to be kept in Cloud Storage as files, batched every few minutes, with no servers or pipelines to manage. What should the data engineer create?
- A.A Cloud Storage subscription on the topic, with file batching by time and sizecorrect
- B.A push subscription to a Cloud Run service that writes each message to an object
- C.A BigQuery subscription followed by a daily export to Cloud Storage
- D.A snapshot of the subscription taken every few minutes
Why: Cloud Storage subscriptions write messages directly into objects in a bucket and batch them into files according to time and size settings, with no extra compute. A Cloud Run service is code to run and creates one object per message. Going through BigQuery and exporting daily adds a step and delay, and subscription snapshots record acknowledgment state rather than writing files.
Open this question on its own page → - Sample · question 6 · Bigtable garbage collection by age
A Bigtable table at Marlowe Weather stores sensor values in a column family that receives a new version of each cell every minute. Only the last 30 days of versions are useful, and storage keeps growing. What should the data engineer configure?
- A.A garbage collection policy on the column family with a maximum age of 30 dayscorrect
- B.A Cloud Scheduler job that deletes old rows through the Bigtable API every night
- C.A lifecycle rule on the table, like the ones used for Cloud Storage buckets
- D.A reduced number of cluster nodes
Why: Garbage collection policies are set per column family and automatically remove cell versions that are older than a maximum age or beyond a maximum number of versions. A scheduled delete job adds code and load. Bigtable has no Cloud Storage-style lifecycle rules, and fewer nodes do not remove data.
Open this question on its own page → - Sample · question 7 · Search indexes for log lookups
Security analysts at Penwortham Bank search a 200 TB BigQuery table of application logs for rare values such as a specific IP address or request ID across many text columns. Each lookup scans most of the table. Which feature makes these needle-in-a-haystack lookups much cheaper and faster?
- A.A materialized view that selects all columns
- B.Clustering the table by every searched column
- C.A search index on the relevant columns, queried with the SEARCH functioncorrect
- D.BI Engine for the project
Why: Search indexes let BigQuery locate the rows that contain specific tokens across text columns, so the SEARCH function reads far less data for point lookups. A materialized view of all columns duplicates the table without helping lookups. Clustering helps with at most four columns used as filters, and BI Engine speeds up aggregations for BI tools rather than token searches.
Open this question on its own page → - Sample · question 8 · Data Studio viewer's credentials
A Data Studio (formerly Looker Studio) report at Thursby Retail uses a BigQuery table protected by row-level security. Each regional manager who opens the report must see only the rows their own BigQuery permissions allow. How should the report's data source be configured?
- A.Set the data source credentials to the owner's credentials
- B.Set the data source credentials to the viewer's credentialscorrect
- C.Use a service account owned by the report creator for the data source
- D.Download the data to a spreadsheet for each manager
Why: With viewer's credentials, each person's own identity is used to query BigQuery, so row-level security filters the data for that viewer. Owner's credentials or a single service account run every query as one identity, showing everyone the same rows. Spreadsheet copies bypass BigQuery's controls and go stale.
Open this question on its own page → - Sample · question 9 · Integer-range partitioning on an ID
Queries on a BigQuery table at Garside Telecom almost always filter on account_id, an integer between 1 and 4,000,000, and the table has no useful date column. The team wants to prune data by account ranges. Which table design allows this?
- A.Ingestion-time partitioning with daily granularity
- B.Time-unit column partitioning on account_id
- C.Creating a separate table for each account
- D.Integer-range partitioning on account_id with a start, end and intervalcorrect
Why: Integer-range partitioning splits a table on an INTEGER column using a start, end and interval, so filters on account_id prune whole partitions. Ingestion time has nothing to do with account_id, and time-unit partitioning requires a DATE, TIMESTAMP or DATETIME column. A table per account would mean millions of tables.
Open this question on its own page → - Sample · question 10 · Approximate distinct counts for dashboards
A dashboard at Wetherby Streaming shows daily unique viewers from a 40 TB BigQuery table, and the exact COUNT(DISTINCT viewer_id) query is slow and expensive. The product team accepts a small statistical error. What should the query use?
- A.COUNT(viewer_id)
- B.APPROX_COUNT_DISTINCT(viewer_id)correct
- C.COUNT(DISTINCT viewer_id) with a LIMIT clause
- D.SUM(1) grouped by viewer_id
Why: APPROX_COUNT_DISTINCT uses a HyperLogLog-based estimate that needs far less memory and computation than an exact distinct count, at the cost of a small error. COUNT counts repeated viewers more than once. LIMIT does not reduce the work of the aggregation, and grouping by viewer_id still processes every distinct value.
Open this question on its own page → - Sample · question 11 · Scheduled queries for a single SQL job
Every morning at 06:00, Hesketh Foods needs one SQL statement to append the previous day's summarized sales into a reporting table in BigQuery. There are no other steps or dependencies. What is the simplest way to automate it?
- A.A Managed Service for Apache Airflow (formerly Cloud Composer) environment with a daily DAG
- B.A BigQuery scheduled query with the destination table set to appendcorrect
- C.A Compute Engine VM running cron and the bq command-line tool
- D.A Dataflow batch pipeline launched by Cloud Scheduler
Why: Scheduled queries run a SQL statement on a schedule through BigQuery Data Transfer Service and can append results to a destination table, with nothing else to manage. An Airflow environment is more than one statement needs and costs money while idle. A cron VM adds a server to maintain, and Dataflow is unnecessary for a single SQL statement.
Open this question on its own page → - Sample · question 12 · Pub/Sub message storage policy regions
A regulation requires that messages published by Aldridge Health to a Pub/Sub topic are stored only in European Union regions, even when publishers connect from elsewhere. What should the data engineer configure?
- A.Message ordering with EU-based ordering keys
- B.A subscription filter on a region attribute
- C.A dead-letter topic created in an EU region
- D.A message storage policy on the topic that allows only EU regionscorrect
Why: A topic's message storage policy lists the regions where Pub/Sub may persist messages, and it can also be enforced through the resource locations organization policy. Ordering keys control delivery order, not storage location. Filters decide which messages a subscription delivers, and a dead-letter topic only receives failed messages.
Open this question on its own page →
Like the sample?
Other practice exams
- AnthropicClaude Certified Architect — Foundations100 questions · $19
- CompTIACompTIA Security+ (SY0-701)100 questions · $19
- ISC2CISSP100 questions · $19
- DatabricksDatabricks Data Engineer Associate100 questions · $19
- DatabricksDatabricks Data Engineer Professional100 questions · $19
- Google CloudGoogle Cloud Associate Cloud Engineer100 questions · $19
- Google CloudGoogle Cloud Professional Cloud Architect100 questions · $19
- HashiCorpTerraform Associate (004)100 questions · $19
- Microsoft Power BI & FabricPower BI Data Analyst (PL-300)100 questions · $19
- Microsoft Power BI & FabricFabric Analytics Engineer (DP-600)100 questions · $19
- SnowflakeSnowPro Core (COF-C03)250 questions · $19
- SnowflakeSnowPro Advanced: Data Engineer100 questions · $19
- SnowflakeSnowPro Advanced: Architect100 questions · $19
- AWSAWS Cloud Practitioner (CLF-C02)100 questions · $19
- AWSAWS Solutions Architect Associate (SAA-C03)100 questions · $19
- AWSAWS AI Practitioner (AIF-C01)100 questions · $19
- Microsoft AzureAzure Fundamentals (AZ-900)100 questions · $19
- Microsoft AzureAzure Administrator (AZ-104)100 questions · $19
- Microsoft AzureAzure AI Fundamentals (AI-901)100 questions · $19
- Microsoft AzureAzure Solutions Architect Expert (AZ-305)100 questions · $19