MLV Tools

Replication and Comparison

Estimate ranking instability caused by variance overlap

Single-run scores can invert true ranking

Load repeated-run metrics, inspect each model distribution, and quantify how often a lower-mean model beats a higher-mean model in one random draw.

MNIST case uses repeated-run metrics generated externally; detection case uses multispectral fusion pipelines from a dedicated utility repository.

Data setup · Models

Simulation setup

Configure trial counts for Monte Carlo and Bootstrap before running the results.

Histogram + fitted normal