Compare component variants under repeated training runs
Sequential ablations can miss the true optimum
See how much each component moves the metric, and whether that change survives repeated training runs.
MNIST case uses repeated-run metrics generated externally; detection case uses multispectral fusion pipelines from a dedicated utility repository.
Data setup · Models
Simulation setup
Configure trial counts for Monte Carlo and Bootstrap before running the results.
Histogram + fitted normal
Results
Model Summary
Model
Best value
N
Mean
Std. dev.
Normal P90
Normality test
Estimate disorder probability over the selected models, averaging N sampled scores per model.
Normality diagnostics justify Monte Carlo representativeness. If diagnostics are weak, Bootstrap is the robust reference.
Shapiro-Wilk
Normality summary
Model
Median
Mean
Skewness
Kurtosis (Fisher)
W
p-value (%)
Mean≈Med
Skew+Kurt
S-W
Monte Carlo samples from fitted normals; Bootstrap samples from empirical data with replacement.
p(!correct order)
N samples/model
Monte Carlo
Bootstrap
Decision path tree
Starting from each fixed condition, this view shows the probability of reaching the global best condition.
ANOVA interaction plot
Each chart shows the mean metric for every level of one factor, with separate lines for each level of the other factor. Non-parallel lines indicate interaction effects between the two factors.
Seed percentile explorer
Each mini-histogram shows the distribution of metric values for one condition. Move the slider to select a seed and see where its result falls in each distribution, with the corresponding percentile.