Your team is conducting an NLP research project to predict authors’ political affiliation from articles they have written. You have a large training dataset structured as follows:
You used the standard 80%-10%-10% distribution for the training, testing, and evaluation subsets. How should the training examples be allocated among the train-test-eval subsets while preserving the 80-10-10 proportion?
A Distribute texts randomly across the train-test-eval subsets: Train set: [TextA1, TextB2, ...] Test set: [TextA2, TextC1, TextD2, ...] Eval set: [TextB1, TextC2, TextD1, ...] B Distribute authors randomly across the train-test-eval subsets: (*) Train set: [TextA1, TextA2, TextD1, TextD2, ...] Test set: [TextB1, TextB2, ...] Eval set: [TexC1,TextC2 ...] C Distribute sentences randomly across the train-test-eval subsets: Train set: [SentenceA11, SentenceA21, SentenceB11, SentenceB21, SentenceC11, SentenceD21 ...] Test set: [SentenceA12, SentenceA22, SentenceB12, SentenceC22, SentenceC12, SentenceD22 ...] Eval set: [SentenceA13, SentenceA23, SentenceB13, SentenceC23, SentenceC13, SentenceD31 ...] D Distribute paragraphs of texts (i.e., chunks of consecutive sentences) across the train-test-eval subsets: Train set: [SentenceA11, SentenceA12, SentenceD11, SentenceD12 ...] Test set: [SentenceA13, SentenceB13, SentenceB21, SentenceD23, SentenceC12, SentenceD13 ...] Eval set: [SentenceA11, SentenceA22, SentenceB13, SentenceD22, SentenceC23, SentenceD11 ...] Show Answer Answer Explanation All material from a given author should remain in a single subset. Otherwise, text, sentences, or paragraphs by the same author can occur in both training and testing/evaluation, allowing author-specific writing patterns to leak across the split and producing an overly optimistic estimate of generalization to unseen authors. Evaluation data must be isolated from training data and represent the real-world examples the model will encounter.
Learn more
Community Discussion