QuestionQ41

Automating and orchestrating ML pipelines

You need to use TensorFlow to train an image-classification model. Your dataset is in a Cloud Storage directory and includes millions of labeled images. Before model training, you need to prepare the data. You want the data-preprocessing and model-training workflow to be as efficient, scalable, and low-maintenance as possible. What should you do?

Explanation

Dataflow is a managed service for batch processing at scale, making it suitable for parallel conversion of millions of labeled images into sharded TFRecord files in Cloud Storage. tf.data.TFRecordDataset reads records from one or more TFRecord files, and managed Vertex AI Training supplies scalable GPU-backed execution without using a notebook instance as the production training environment.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!