Computer vision
Vision Workbench
An image-classification experiment with a closer look at the examples a model gets wrong.
The link opens my GitHub profile until a project repository is added.
The problem
A single accuracy score can hide repeated failures on a particular class or capture condition. This concept puts error analysis next to the training workflow so a model can be inspected beyond its headline score.
How it would work
Use a public image dataset and preserve its documented split. Audit duplicates and label mappings before applying any augmentation.
Compare a frozen pretrained backbone with a small fine-tuning run under the same data and evaluation conditions. Keep training augmentations separate from deterministic evaluation transforms.
Build an error gallery with true labels, predictions, and model scores. Group confusions by class, and inspect whether backgrounds or repeated images explain apparently strong performance.
What to evaluate
A proposed evaluation plan for this example:
- Report per-class precision and recall alongside overall accuracy.
- Inspect a confusion matrix and a reproducible sample of incorrect predictions.
- Select settings on validation data, then evaluate the chosen configuration once on the reserved test set.