Evaluate a brain-tumor segmentation model with more than one score: report per-region Dice for overlap, HD95 (or another fully specified surface-distance measure) for boundary error, and sensitivity/specificity or lesion-wise detection measures for misses and false positives. Calculate results per case before summarizing the cohort, and document label definitions, image spacing, metric implementation, empty-mask handling, and aggregation. These are research-evaluation measures, not proof that a segmentation is clinically acceptable.
Why no single score is enough
Segmentation quality has several dimensions. A model can overlap much of a reference mask yet place part of its boundary far away; it can also have a reasonable average score while missing a small region or an entire lesion in an individual case. Overlap, contour distance, and detection measures answer different questions, so report them together rather than treating one as a complete verdict.
- Overlap: How much of the predicted volume agrees with the reference?
- Boundary error: How far apart are the predicted and reference contours?
- Misses and excess: How much reference-positive tissue was missed, and how much negative tissue was labeled positive?
- Clinical perception: Do qualified reviewers consider the result usable for the intended task?
What each metric tells you
Dice measures overlap
The Dice similarity coefficient (DSC) compares the intersection of predicted and reference masks with their combined size. Under the common binary definition, it ranges from no overlap to perfect overlap. Dice is intuitive, but it does not express how far a displaced contour lies from the reference. The same amount of missed or added volume can be arranged close to the correct boundary or far from it. A modest number of voxels can also change Dice substantially for a small target. The 2022 review of brain-tumor imaging discusses Dice alongside other voxel- and surface-based measures: Preoperative Brain Tumor Imaging: Models and Software for Segmentation and Standardized Reporting.
Hausdorff distance exposes boundary outliers
Hausdorff distance measures the worst nearest-point separation between two boundaries, considering distances in both directions. A small remote false-positive island or one extreme mismatch can dominate the result. The historical BRATS benchmark documented this outlier sensitivity and used the 95th percentile of surface distances as a more robust alternative. In one example from that older benchmark, a method missed all active-tumor voxels in three volumes: its average Dice still looked favorable, while mean Hausdorff distance was dominated by those failures. This illustrates a metric limitation; it is not a current model comparison. See Menze et al., The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS).
#1 Best Overall
HD95 reduces, but does not eliminate, outlier sensitivity
HD95 is the 95th percentile of surface distances under a specified implementation. It reduces the influence of the most extreme tail compared with maximum Hausdorff distance, but it is not outlier-proof. Implementations can differ in surface extraction, whether and how directions are combined, percentile conventions, and use of image spacing. Report the exact convention and implementation. When image geometry is available, express distance in millimetres; a value in voxels is not directly comparable across images with different spacing.
The official BraTS 2019 evaluation page specifies Dice and Hausdorff distance (95%) for that challenge’s segmentation task: BraTS 2019 evaluation. That historical specification is an example, not a guarantee that another dataset or current challenge uses the same protocol.
Rank #2
Complementary surface and detection measures
Average symmetric surface distance (ASSD) summarizes typical bidirectional contour separation, complementing HD95’s focus on the distance tail. Surface Dice, also called a normalized surface-distance measure in some contexts, reports how much of a surface falls within a specified tolerance. Boundary F1 likewise evaluates boundary precision and recall under a tolerance. These measures are not interchangeable: state the tolerance and units, and explain why the chosen measure suits the task. The 2022 review and a 2024 radiotherapy auto-segmentation review describe complementary boundary-measure approaches: 2022 brain-tumor imaging review and 2024 NRG Oncology assessment.
Sensitivity (recall) is the fraction of reference-positive voxels recovered; specificity is the fraction of reference-negative voxels correctly rejected. These help reveal under- and over-segmentation alongside Dice. For multifocal tumors or tasks where every lesion matters, add lesion-wise detection counts or precision and recall. Whole-volume voxel metrics can hide a missed lesion, so distinguish voxel-wise, patient-wise, and instance-wise reporting when relevant. The BraTS 2019 page documents Dice, HD95, sensitivity, and specificity in its evaluation scheme: BraTS 2019 evaluation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Define regions before comparing models
For the BraTS 2020 adult glioma task, the benchmark defines enhancing tumor (ET), tumor core (TC), and whole tumor (WT). TC comprises ET plus necrotic and non-enhancing core components; WT comprises TC plus peritumoral edema. These are benchmark-specific label conventions, not universal definitions for other tumor types, datasets, treatment stages, or annotation protocols. The official task definitions are at BraTS 2020 tasks.
Report scores separately for every relevant region. A composite or macro-average can be useful if its calculation is explicit, but it should not replace visible per-region results: a larger region can conceal poor performance on a smaller subregion. Present case-level distributions as well as cohort summaries, including failures and outliers where appropriate. There is no single aggregation method prescribed for every study.
A reproducible evaluation workflow
- Specify the task and target. Name the tumor population and imaging setting, define each label and subregion, describe how reference annotations were created, and say whether the task is whole-volume semantic segmentation or lesion-wise detection.
- Freeze the evaluation protocol. Use held-out cases that were not used to tune thresholds or select the model. Record preprocessing, postprocessing, image geometry and voxel spacing, label mapping, and how empty masks or missing labels are handled. Challenge configurations are task-specific, so do not assume another dataset shares them.
- Calculate complementary measures per region and case. At minimum, report Dice and HD95. Add sensitivity/specificity or lesion-wise measures when missed lesions and false positives matter. Add ASSD or a tolerance-based surface measure when typical contour agreement is important, stating tolerance and units.
- Summarize without hiding failures. State the number of cases and the aggregation rule. Show case-level distributions and meaningful outliers alongside means; do not rely on a single cohort average. Metric choice can change rankings, and the historical BRATS benchmark documents an example where aggregate Dice obscured severe failures (benchmark article).
- Compare models on paired cases. Evaluate systems on the same test cases, describe uncertainty and the statistical comparison used, and avoid claiming a meaningful improvement from a tiny score change without an appropriate analysis. The appropriate test depends on study design and outcome distribution; the cited sources do not establish one universal choice.
- Add expert review when perceived quality matters. Define reviewer qualifications, rubric, blinding, and how disagreement is handled. In a 2023 RSNA study, only 2.8% (five of 180 surveyed articles) included clinical-expert evaluation of segmentation quality. In that study’s experiment, expert-rating interrater agreement was Krippendorff α = 0.34, and correlation between Dice and mean expert quality rating was Kendall tau = 0.23. These are findings from that surveyed literature and experiment, not prevalence or agreement estimates for all medical AI research. The authors reported that quality ratings varied with ambiguity in tumor boundaries and perceptual differences, and that existing metrics did not capture clinical perception: RSNA 2023 expert-centered evaluation.
- Record software and versions. The BraTS Evaluation repository describes a Python package that accepts reference and prediction NIfTI files, provides task configurations, and can produce JSON summaries and CSV reports. Its documentation also describes instance-wise HD95 and normalized surface-distance capabilities. Verify that the package version and configuration match the dataset, then report both: BraTS Evaluation repository.
How to compare two models fairly
| Comparison question | Useful evidence | What to check |
|---|---|---|
| Does overlap improve? | Per-region Dice | Are gains consistent across regions, or driven by a larger region? |
| Are severe contour errors present? | HD95 | Are there high-distance cases or isolated failures? |
| How far apart are contours typically? | ASSD or tolerance-based surface measure | Are units and any surface tolerance justified and stated? |
| Are tissue or lesions missed, or extra regions labeled? | Sensitivity, specificity, precision/recall, and lesion-wise measures where relevant | Do voxel-level and lesion-level results tell different stories? |
| Is performance reliable and reproducible? | Case-level distributions, qualified expert review, and protocol details | Were the same cases, labels, geometry, implementation, empty-mask rules, and aggregation used? |
Limits of metric-based conclusions
No universal clinically acceptable Dice or HD95 threshold is established by the cited sources. Acceptability depends on the target, intended use, reference labels, image resolution, annotation uncertainty, and consequences of error. A 2015 BRATS benchmark reported 74%–85% Dice inter-rater agreement for human raters segmenting tumor subregions in that benchmark. Treat that as evidence of difficulty in that particular annotation task, not as a target score or a universal range for human agreement (Menze et al.).
Expert assessment is useful when clinical perception matters, but it is not infallible: reviewers can disagree, so report the rubric and agreement rather than presenting review as an unquestionable gold standard. Metric results support a defined research comparison; on their own, they do not establish clinical readiness or diagnostic suitability.
Recommended Free Tools
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

