← All writing
ML fundamentalsSeptember 20262 min read

A validation score you can trust

Before improving a model’s score, check what the score is measuring. A careful split often tells you more than a more complicated model.

On this page

Decide what “unseen” means

Imagine classifying product photos. If near-identical photographs of one product appear in both training and validation, the score may measure familiarity with that product rather than generalization to a new one. The rows are different; the underlying example is almost the same.

Choose the split around the way the model will be used. Group related observations when testing on new entities. Preserve temporal order when predicting future observations. A random split can be appropriate for independent examples, but it does not automatically fit every prediction problem.

Each split has a different job. The proportions shown are illustrative, not a recommended ratio; the right split depends on the data and task.

Preprocessing can leak information

Leakage can happen before the model sees a label. For example, a scaler fitted on the full dataset has already used information from the held-out observations. Split first, fit preprocessing on the training partition, and use the learned transformation on the other partitions.

Keep fitting inside the training splitpython
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

X, y = load_iris(return_X_y=True)
X_train, X_valid, y_train, y_valid = train_test_split(
    X, y, test_size=0.25, stratify=y, random_state=7
)
model = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=500),
)
model.fit(X_train, y_train)
print(model.score(X_valid, y_valid))

The pipeline fits its scaler using X_train. Its score method transforms X_valid using that same fitted scaler before evaluating the classifier. This small example uses a validation split for demonstration; a final test set would be reserved separately.

Validation is feedback

Use validation data to compare model settings. With cross-validation, preprocessing should be fitted separately inside each fold, which is another reason to include it in a pipeline. A group-aware or time-aware splitter can preserve the structure of the prediction problem.

Inspect the errors, too

Suppose the overall score improves while one small class gets worse. That can be a meaningful tradeoff, but it should be visible. Read a few correct and incorrect examples, check the label mapping, and record which classes or capture conditions deserve another look.

  • Keep the split definition and random seed with the experiment.
  • Compare models on the same held-out examples.
  • Check duplicates and unavailable-at-prediction-time features.
  • Write down a failure case alongside the headline metric.

Sources & further reading

Primary references for the concepts and APIs in this article.

  1. scikit-learn — Data leakage and common pitfalls (opens in a new tab)
  2. scikit-learn — Cross-validation (opens in a new tab)