QuestionQ59

Scaling prototypes into ML models

You are training an object-detection model on a Cloud TPU v2, but training is taking longer than expected. Based on this simplified trace from a Cloud TPU profile, what action should you take to reduce training time cost-effectively?

Question Image

Explanation

The workload is input-bound: data extraction, transformation, and network loading occur before each TPU training interval, leaving the TPU idle while it waits for the next batch. Parallel file reads and preprocessing, together with prefetching, overlap input preparation with training and keep the TPU supplied with data without requiring more expensive accelerator hardware. Google Cloud recommends parallel reads, parallel map() calls, and prefetch() to mitigate TPU input-processing bottlenecks.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!