A validation score you can trust
Before improving a model’s score, check what the score is measuring. A careful split often tells you more than a more complicated model.
On this page
Decide what “unseen” means
Imagine classifying product photos. If near-identical photographs of one product appear in both training and validation, the score may measure familiarity with that product rather than generalization to a new one. The rows are different; the underlying example is almost the same.
Choose the split around the way the model will be used. Group related observations when testing on new entities. Preserve temporal order when predicting future observations. A random split can be appropriate for independent examples, but it does not automatically fit every prediction problem.
Group related examples or respect time order when the task requires it.
Preprocessing can leak information
Leakage can happen before the model sees a label. For example, a scaler fitted on the full dataset has already used information from the held-out observations. Split first, fit preprocessing on the training partition, and use the learned transformation on the other partitions.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
X, y = load_iris(return_X_y=True)
X_train, X_valid, y_train, y_valid = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=7
)
model = make_pipeline(
StandardScaler(),
LogisticRegression(max_iter=500),
)
model.fit(X_train, y_train)
print(model.score(X_valid, y_valid))The pipeline fits its scaler using X_train. Its score method transforms X_valid using that same fitted scaler before evaluating the classifier. This small example uses a validation split for demonstration; a final test set would be reserved separately.
Validation is feedback
Use validation data to compare model settings. With cross-validation, preprocessing should be fitted separately inside each fold, which is another reason to include it in a pipeline. A group-aware or time-aware splitter can preserve the structure of the prediction problem.
Inspect the errors, too
Suppose the overall score improves while one small class gets worse. That can be a meaningful tradeoff, but it should be visible. Read a few correct and incorrect examples, check the label mapping, and record which classes or capture conditions deserve another look.
- Keep the split definition and random seed with the experiment.
- Compare models on the same held-out examples.
- Check duplicates and unavailable-at-prediction-time features.
- Write down a failure case alongside the headline metric.
Sources & further reading
Primary references for the concepts and APIs in this article.