Mode
Text Size
Log in / Sign up

Clinical Segmentation Evaluation Scale shows strong correlation with IoU for AI nerve segmentationNew scale helps experts judge accuracy of nerve imaging AI

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that C-SES shows strong correlation with IoU for nerve segmentation, but results are preliminary and hypothesis-generating.

This guideline provides a pilot validation of the Clinical Segmentation Evaluation Scale (C-SES) to assess AI-based nerve segmentation on ultrasound. The evaluation focused on content validity, concurrent criterion validity, and reliability compared to Intersection over Union (IoU) metrics across 74 ultrasound sequences.

The study found a strong correlation between C-SES and IoU (Pearson's r = 0.861). Reliability measures were high, with an intra-rater ICC(2,1) of 0.905 and inter-rater ICC values of 0.869 for aggregated data and 0.940 for individual raters. Youden thresholds for non-expert usefulness were identified as 3.51 for C-SES and 0.206 for IoU, while novice usefulness thresholds were 4.689 for C-SES and 0.299 for IoU.

The authors note that these proposed thresholds are preliminary and hypothesis-generating rather than universal clinical benchmarks. Further validation is required before the scale can be generalized beyond expert assessors and the specific sample studied. The C-SES appears to estimate objective segmentation metrics, but its utility as a standard tool requires further evidence.

When doctors use artificial intelligence to map out nerves on ultrasound scans, they need a consistent way to measure how well the technology is actually performing. Currently, it can be difficult to get a clear picture of accuracy across different tools and experts.

A new scoring system called C-SES was tested to solve this problem. Researchers compared this scale against standard technical metrics used to measure overlap in images. They found that the C-SES score had a strong correlation with these technical measurements, and it proved highly reliable when used by multiple experts to grade 74 different ultrasound sequences.

While the results are promising, it is important to note that this was a pilot study. The specific scoring thresholds identified are preliminary and not yet established as universal clinical benchmarks. More research is needed to confirm these findings before the tool can be used broadly outside of expert settings.

What this means for you:
A new scale provides a reliable way for experts to measure how accurately AI identifies nerves in ultrasound scans.

Common questions

How reliable is this new scoring system?

The study found that the C-SES scale has excellent reliability when used by a single person and high reliability when results are combined across multiple experts. These findings suggest it is a consistent tool for expert raters to evaluate AI performance.

Is this system ready for everyday clinical use?

Not quite yet. Because this was a pilot study, the results are preliminary and intended to generate new ideas rather than serve as universal benchmarks. More validation is needed before it can be used outside of expert settings.

How does this help with AI-based nerve imaging?

The C-SES scale helps estimate objective metrics for how well AI identifies nerves on ultrasound scans. It showed a strong correlation with standard technical measures, making it easier for experts to judge the quality of the technology.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedAug 2026
View Original Abstract ↓
BackgroundThe objective metrics for AI-based nerve segmentation are time-consuming and may not fully reflect their clinical usefulness. We aimed to 1) conduct a pilot validation study of the Clinical Segmentation Evaluation Scale (C-SES), as an estimation for an objective metric, and 2) establish a threshold on those scales, because the thresholds were derived for both the new C-SES scale AND for objective metrics (IoU, DSC) to determine the usefulness of AI predictions.MethodsSeven experts rated 74 ultrasound sequences twice using AI-based nerve recognition software (cNerve) in 3 anatomical regions. In Step 1, the C-SES was evaluated for 1) content validity (expert appraisal of relevance, comprehensiveness, and clarity), 2) concurrent criterion validity (by correlating the C-SES with a reference standard, the Intersection over Union—IoU), and 3) reliability (intra and inter-rater) and error measurement. In an additional Step 2, expert-perceived usefulness for non-experts and novice anesthesiologists was evaluated to derive candidate thresholds for Intersection over Union (IoU) and C-SES metrics within the investigated clinical context.ResultsIn Step 1, content validity was preliminarily supported by expert consensus. Concurrent criterion validity was strong (Pearson’s r = 0.861). Reliability was excellent within raters [mean ICC(2,1) = 0.905] and high for aggregated inter-rater ratings [ICC(2,3) = 0.869; ICC(2,k) = 0.940]. Precision increased when several raters were pooled. In Step 2, the Youden thresholds for non-expert usefulness were 0.206 for the IoU and 3.51 for the C-SES; for novice usefulness, the thresholds were 0.299 and 4.689, respectively.ConclusionIn this pilot validation, the C-SES appeared to estimate objective segmentation metrics, although further validation is needed. The proposed thresholds are preliminary and hypothesis-generating rather than universal clinical benchmarks and require confirmation before generalization beyond expert assessors and this spectrum-enriched sample.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.