MLV Tools
Estimate ranking instability caused by variance overlap
MNIST case uses repeated-run metrics generated externally; detection case uses multispectral fusion pipelines from a dedicated utility repository.
Configure trial counts for Monte Carlo and Bootstrap before running the results.
Estimate disorder probability over the selected models, averaging N sampled scores per model.
Normality diagnostics justify Monte Carlo representativeness. If diagnostics are weak, Bootstrap is the robust reference.
Monte Carlo samples from fitted normals; Bootstrap samples from empirical data with replacement.