QuestionQ1

Scaling prototypes into ML models

You are pre-training a large language model on Google Cloud. The model contains custom TensorFlow operations in its training loop. Training will use a large batch size, and you anticipate that it will run for several weeks. You need to configure a training architecture that minimizes both training duration and compute costs. What should you do?

Explanation

Custom TensorFlow operations require a GPU-based training design unless the operations are specifically available and compatible with Cloud TPU; Cloud TPU supports only its documented set of TensorFlow APIs and graph operators. tf.distribute.MultiWorkerMirroredStrategy supports synchronous training across multiple workers with multiple GPUs. Eight a2-megagpu-16g workers provide 128 A100 GPUs while using half as many VM hosts as sixteen a2-highgpu-8g workers, reducing duplicated host resources while retaining the same aggregate GPU count.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!