MEDOTTER · BENCHMARKS · LEADERBOARD
Rank a model.
See where it stands.
Approved Hugging Face models are evaluated with the MedOtter harness and published here with versioned context. Public submissions are coming soon.
| Rank | Model | Details | Testing set | |||
|---|---|---|---|---|---|---|
| DATASET-AVG | CASE-AVG | |||||
| Rank 1 | MedOtter Internal Alpha Zero BaselineZero baseline MedOtter | 0.000 | 373.130 | 0.000 | 373.130 | |
| Rank 2 | Swin-UNetunpinned Angelou0516 | 0.707 | 56.814 | 0.707 | 60.716 | |
| Rank 3 | VM-UNetunpinned Angelou0516 | 0.770 | 50.707 | 0.770 | 54.855 | |
| 04 | zero-baselineunpinned | 0.000 | 373.130 | 0.000 | 373.130 | |
| 05 | Attention U-Netunpinned Angelou0516 | 0.723 | 62.297 | 0.718 | 67.668 | |
| 06 | RWKV-UNetunpinned Angelou0516 | 0.770 best in column | 49.589 | 0.769 | 54.501 | |
| 07 | monai-segresnetCNN | 0.708 | 29.953 best in column | 0.802 best in column | 19.501 best in column | |
| 08 | TransUNetunpinned Angelou0516 | 0.743 | 58.041 | 0.742 | 63.040 | |
Scores are each task's metrics on its headline region, on the frozen test split. DATASET-AVG averages per dataset; CASE-AVG weights by case count (default). Lower is better for distance metrics (e.g. HD95); higher is better for overlap and detection metrics. A ★ marks the best value in each column. A “stale” dataset label means the newest available score predates the task's current dataset version and is shown only as a fallback while re-evaluation is pending. Empty-structure convention: both empty → Dice 1.0 / HD95 0 mm; one side empty → Dice 0 / HD95 373.13 mm (fixed penalty).