MLV Tools

Standardizing Comparisons

Align metrics and protocols before comparing models

A single reported result is often one lucky run

See how many repeated runs it takes for scores and rankings to settle, so a comparison stops hinging on one lucky run.

MNIST case uses repeated-run metrics generated externally; detection case uses multispectral fusion pipelines from a dedicated utility repository.

Data setup · Models

Simulation setup

Configure trial counts for Monte Carlo and Bootstrap before running the results.

Survival function plot