You are pre-training a large language model on Google Cloud. The model contains custom TensorFlow operations in its training loop. Training will use a large batch size, and you anticipate that it will run for several weeks. You need to configure a training architecture that minimizes both training duration and compute costs. What should you do?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion