QuestionQ18

Automating and orchestrating ML pipelines

Your team is conducting an NLP research project to predict authors’ political affiliation from articles they have written. You have a large training dataset structured as follows:

Question Image

You used the standard 80%-10%-10% distribution for the training, testing, and evaluation subsets. How should the training examples be allocated among the train-test-eval subsets while preserving the 80-10-10 proportion?

Explanation

All material from a given author should remain in a single subset. Otherwise, text, sentences, or paragraphs by the same author can occur in both training and testing/evaluation, allowing author-specific writing patterns to leak across the split and producing an overly optimistic estimate of generalization to unseen authors. Evaluation data must be isolated from training data and represent the real-world examples the model will encounter.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!