MEDOTTER · BENCHMARKS · LEADERBOARD

Rank a model.
See where it stands.

Approved Hugging Face models are evaluated with the MedOtter harness and published here with versioned context. Public submissions are coming soon.

8
TASKS
8
MODELS
4
DATASETS
2
METRICS
MedOtter model leaderboard
RankModelDetailsTesting set
DATASET-AVGCASE-AVG
Rank 1
MedOtter Internal Alpha Zero BaselineZero baseline
MedOtter
0.000373.1300.000373.130
Rank 2
Swin-UNetunpinned
Angelou0516
0.70756.8140.70760.716
Rank 3
VM-UNetunpinned
Angelou0516
0.77050.7070.77054.855
04
0.000373.1300.000373.130
05
Angelou0516
0.72362.2970.71867.668
06
RWKV-UNetunpinned
Angelou0516
0.770 best in column49.5890.76954.501
07
monai-segresnetCNN
0.70829.953 best in column0.802 best in column19.501 best in column
08
TransUNetunpinned
Angelou0516
0.74358.0410.74263.040

Scores are each task's metrics on its headline region, on the frozen test split. DATASET-AVG averages per dataset; CASE-AVG weights by case count (default). Lower is better for distance metrics (e.g. HD95); higher is better for overlap and detection metrics. A ★ marks the best value in each column. A “stale” dataset label means the newest available score predates the task's current dataset version and is shown only as a fallback while re-evaluation is pending. Empty-structure convention: both empty → Dice 1.0 / HD95 0 mm; one side empty → Dice 0 / HD95 373.13 mm (fixed penalty).