QuestionQ54

Scaling prototypes into ML models

You are building an image-recognition model in PyTorch using the ResNet50 architecture. Your code runs successfully on a small subsample on your local laptop. The complete dataset contains 200k labeled images. You need to scale the training workload quickly while keeping costs as low as possible, and you plan to use 4 V100 GPUs. What should you do?

Explanation

Vertex AI custom training can run a packaged Python training application in a prebuilt PyTorch training container and attach the required GPU resources. This provides managed, scalable GPU training without the operational work of creating and maintaining a Kubernetes cluster or manually configuring a Compute Engine training VM.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!