About the Exam

The exam is for experienced data engineers working on Google Cloud. It covers designing data processing systems, ingesting and processing data, storing data, preparing data for analysis, and maintaining automated data workloads. Passing demonstrates you can build and operate secure, scalable data infrastructure on Google Cloud.

Exam Topics

  • Designing data processing systems22%
  • Ingesting and processing the data25%
  • Storing the data20%
  • Preparing and using data for analysis15%
  • Maintaining and automating data workloads18%

How to Use This Practice Exam

  1. Browse — Read each question, select your answer, and reveal the explanation.
  2. Exam Mode — Simulate real exam conditions with a timed session and score report.
  3. Learn Mode — Spaced repetition schedules questions you struggle with for long-term retention.

Download the Full Exam PDF

Get every question and answer in a clean, printable PDF built for offline study. Purchase once, keep permanent access, and re-download the latest version anytime.

Last updated July 2, 2026 at 1:35 PM

Topic filter
Retired questions
Question sort
Questions per page

QuestionQ1

Ingesting and processing the data

You analyze user clickstream data to personalize content recommendations. The data arrives continuously and must be processed with low latency, including transformations such as sessionization (grouping clicks by user within a time window) and aggregation of user activity. You need to identify a scalable solution that can handle millions of events per second and remain resilient to late-arriving data. What should you do?

Explanation

Pub/Sub provides scalable, reliable ingestion for continuous event streams. Dataflow with Apache Beam performs low-latency stateful stream processing, including session windows and aggregations, and uses event-time watermarks, triggers, and allowed-lateness settings to handle out-of-order and late data. BigQuery is a suitable analytical store for the processed clickstream data.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!

QuestionQ2

Ingesting and processing the data

A web server publishes click events as messages to a Pub/Sub topic. The server includes an eventTimestamp attribute in each message that records when the click occurred. You have a Dataflow streaming job that reads this Pub/Sub topic through a subscription, performs some transformations, and writes the results to another Pub/Sub topic for the advertising department.

The advertising department must receive every message within 30 seconds of its corresponding click, but reports that messages arrive late. Your Dataflow job has system lag of about 5 seconds and data freshness of about 40 seconds. Inspection of several messages shows no more than 1 second of lag between eventTimestamp and publishTime. What is the issue, and what should you do?

Explanation

Data freshness measures the difference between an element’s event time and processing time; a value of about 40 seconds shows that some elements are being processed too late to meet a 30-second requirement. The approximately 1-second delay before Pub/Sub publication does not account for the lateness, so the pipeline requires performance optimization or additional worker capacity.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!

QuestionQ3

Designing data processing systems

You run a logistics company and want to make event delivery from vehicle-based sensors more reliable. You operate small data centers worldwide to capture these events, but the leased lines connecting your event-collection infrastructure to your event-processing infrastructure are unreliable and have unpredictable latency. You want to resolve this in the most cost-effective manner. What should you do?

Explanation

Cloud Pub/Sub provides managed, asynchronous event ingestion that decouples publishers from downstream processing and is designed for scalable, reliable event delivery. Publishing directly from the data-acquisition devices removes the leased-line dependency without the operational cost of self-managed Kafka clusters or the expense of Cloud Interconnect. Pub/Sub is also a documented fit for streaming data from IoT devices and sensors.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!

QuestionQ4

Ingesting and processing the data

You have a BigQuery table that receives data directly from a Pub/Sub subscription. The ingested data is encrypted using a Google-managed encryption key. You must comply with a new organization policy requiring keys from a centralized Cloud Key Management Service (Cloud KMS) project to encrypt data at rest. What should you do?

Explanation

BigQuery supports customer-managed encryption keys (CMEK) from Cloud KMS for data at rest. Creating a CMEK-protected BigQuery table and migrating the old table’s data ensures the stored historical and destination data is protected by the centrally managed key; changing only the ingestion path or Pub/Sub topic does not re-encrypt the BigQuery table’s existing data.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!

QuestionQ5

Preparing and using data for analysis

You monitor and optimize your team’s BigQuery instance. A particular daily report that uses a large JOIN operation is consistently slow. You want to inspect the query’s execution plan to identify potential performance bottlenecks within the JOIN as quickly as possible. What should you do?

Explanation

BigQuery’s query execution graph presents the execution plan as stages and exposes timing, resource-use, and step details. A JOIN stage can be inspected directly to identify stages that dominate execution time or resource consumption, making it the most direct way to diagnose JOIN-related bottlenecks.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!
Know a question that should be here? Contribute to this exam
Back home