MLV Tools
Align metrics and protocols before comparing models
MNIST case uses repeated-run metrics generated externally; detection case uses multispectral fusion pipelines from a dedicated utility repository.
Configure trial counts for Monte Carlo and Bootstrap before running the results.
| Model | Best value | N | Mean | Std. dev. | Normal P90 |
|---|
Estimate disorder probability over the selected models, averaging N sampled scores per model.
Normality diagnostics justify Monte Carlo representativeness. If diagnostics are weak, Bootstrap is the robust reference.
| Shapiro-Wilk | Normality summary | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Median | Mean | Skewness | Kurtosis (Fisher) | W | p-value (%) | Mean≈Med | Skew+Kurt | S-W |
For the selected percentile P, probability that at least one of N samples exceeds the fitted-normal threshold. One line per model + average.
Each line is the model-average exceedance curve for one percentile in the computed range. Shows how the curve shifts as the threshold tightens.
The sampling error (standard error of the mean) decreases as σ/√N. A few extra training runs can substantially reduce estimation uncertainty. The derivative plot below highlights where diminishing returns set in (Section 4.2).