QuestionQ21

Scaling prototypes into ML models

You are building an ML model that predicts house prices. During data preparation, you find that an important predictor variable—the distance to the nearest school—is frequently missing and has low variance. Every instance (row) in the data is important. How should the missing data be handled?

Explanation

Regression imputation preserves the important rows while estimating missing numeric distances from relationships in the available data. Deleting rows would discard required instances, and assigning zero would introduce an arbitrary value that can bias the model. Google’s machine-learning guidance specifically identifies a regression model trained on existing feature data as a method for imputing missing numerical data.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!