HUGIML OpenML-CC18 Benchmark Dashboard

Official OpenML train/test indices, fold-local model selection, predictive performance, and model-inspection evidence.

Overall model comparison

Performance, ranks, inspection efficiency, and runtime

Performance, ranks & runtime

HUGIML inspection efficiency — baseline / HUGIML

Ratios are baseline model inspection units divided by HUGIML model inspection units, on tasks where HUGIML AUC is within tolerance.

Model inspection units versus AUC

One point per completed task/model pair

The x-axis is logarithmic when every displayed inspection value is positive.

Dataset explorer

Inspect paired results for one OpenML task

Dataset-level metric comparison

Three measures × all models

Each row compares ROC-AUC, balanced accuracy, and F1 for the same OpenML task. Bold marks the best score; italics mark the second-best score. Ranking is based on the displayed four-decimal values, so displayed ties receive identical styling.

HUGIML paired tests

Wilcoxon tests with Holm adjustment

Evaluation protocol

Run configuration

Model search spaces

Exact configuration recorded with this run
Show model grids