QuestionQ1
Scaling prototypes into ML modelsYou are pre-training a large language model on Google Cloud. The model contains custom TensorFlow operations in its training loop. Training will use a large batch size, and you anticipate that it will run for several weeks. You need to configure a training architecture that minimizes both training duration and compute costs. What should you do?
QuestionQ2
Serving and scaling modelsYou built a custom model that carries out several memory-intensive preprocessing tasks before making a prediction. You deployed the model to a Vertex AI endpoint and confirmed that results were returned within a reasonable time. After directing user traffic to the endpoint, you find that it does not autoscale as expected when it receives multiple requests. What should you do?
Community Discussion
QuestionQ3
Architecting low-code AI solutionsYou need to create classification workflows across several structured datasets that are currently stored in BigQuery. Because you will perform the classification multiple times, you want to complete these steps without writing code: exploratory data analysis, feature selection, model building, training, hyperparameter tuning, and serving. What should you do?
Community Discussion
QuestionQ4
Monitoring AI solutionsYou work for a magazine distributor and must build a model that predicts which customers will renew their subscriptions in the coming year. Using your company’s historical data as the training set, you created a TensorFlow model and deployed it to AI Platform. You need to identify which customer attribute has the greatest predictive influence for each prediction served by the model. What should you do?
Community Discussion
QuestionQ5
Serving and scaling modelsYou work for an online travel agency that also sells advertising placements on its website to other companies. You have been asked to predict the most relevant web banner that a user should see next. Security is important to your company. The model-latency requirement is 300ms@p99, the inventory contains thousands of web banners, and your exploratory analysis has shown that navigation context is a good predictor. You want to implement the simplest solution. How should you configure the prediction pipeline?

Community Discussion