QuestionQ5

Serving and scaling models

You work for an online travel agency that also sells advertising placements on its website to other companies. You have been asked to predict the most relevant web banner that a user should see next. Security is important to your company. The model-latency requirement is 300ms@p99, the inventory contains thousands of web banners, and your exploratory analysis has shown that navigation context is a good predictor. You want to implement the simplest solution. How should you configure the prediction pipeline?

Explanation

Cloud Bigtable provides scalable, low-latency, high-throughput keyed reads and writes, which suits storing and retrieving each user's navigation context on the online prediction path. An App Engine gateway provides a controlled application layer between the website client and backend services, while managed AI Platform Prediction avoids the additional operational work of hosting the model on Google Kubernetes Engine.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!