torch_em.data.datasets.medical.segrap2025

The SegRap2025 dataset contains annotations for lymph node clinical target volume (LN CTV) segmentation in head and neck CT scans of nasopharyngeal carcinoma patients.

It is the "Task02: LN CTV Segmentation" part of the SegRap2025 challenge (https://hilab-git.github.io/SegRap2025_Challenge) and is unrelated to the organs-at-risk data of SegRap2023 covered in 'segrap.py': it comprises a different set of 440 CT scans from 262 patients, collected across 5 cohorts (4 centers), each with a pre-aligned pair of a non-contrast and a contrast-enhanced CT scan. 120 patients from the internal cohort provide the official training split (labelled), the remaining 4 testing cohorts (142 patients) are distributed with labels as part of this public release too, and are exposed here as the 'test' split.

NOTE: The label legend is not documented in the released data. Verified on the data: the label volumes contain the ids 0 to 6, i.e. background and 6 disjoint LN CTV sub-levels.

The dataset is a redistribution of the official SegRap2025 Task02 training data (https://doi.org/10.6084/m9.figshare.26793622, CC BY 4.0) via Figshare. The archive is password protected, the password ('lnctvseg@uestc') is published on the challenge dataset page.

This dataset is from the publication https://doi.org/10.48550/arXiv.2601.20575. Please cite it if you use this dataset in your research.

  1"""The SegRap2025 dataset contains annotations for lymph node clinical target volume (LN CTV) segmentation
  2in head and neck CT scans of nasopharyngeal carcinoma patients.
  3
  4It is the "Task02: LN CTV Segmentation" part of the SegRap2025 challenge
  5(https://hilab-git.github.io/SegRap2025_Challenge) and is unrelated to the organs-at-risk data of SegRap2023
  6covered in 'segrap.py': it comprises a different set of
  7440 CT scans from 262 patients, collected across 5 cohorts (4 centers), each with a pre-aligned pair of a
  8non-contrast and a contrast-enhanced CT scan. 120 patients from the internal cohort provide the official training
  9split (labelled), the remaining 4 testing cohorts (142 patients) are distributed with labels as part of this
 10public release too, and are exposed here as the 'test' split.
 11
 12NOTE: The label legend is not documented in the released data. Verified on the data: the label volumes contain
 13the ids 0 to 6, i.e. background and 6 disjoint LN CTV sub-levels.
 14
 15The dataset is a redistribution of the official SegRap2025 Task02 training data
 16(https://doi.org/10.6084/m9.figshare.26793622, CC BY 4.0) via Figshare. The archive is password protected, the
 17password ('lnctvseg@uestc') is published on the challenge dataset page.
 18
 19This dataset is from the publication https://doi.org/10.48550/arXiv.2601.20575.
 20Please cite it if you use this dataset in your research.
 21"""
 22
 23import os
 24import zipfile
 25from glob import glob
 26from shutil import which
 27from subprocess import run
 28from natsort import natsorted
 29from typing import Union, Tuple, Literal, List
 30
 31from torch.utils.data import Dataset, DataLoader
 32
 33import torch_em
 34
 35from .. import util
 36
 37
 38URL = "https://ndownloader.figshare.com/files/48684664"
 39CHECKSUM = "995215f98844f1a5718a1b963fe5615dc3723d506a889503617fcd99aea052a1"
 40
 41PASSWORD = "lnctvseg@uestc"
 42
 43MODALITIES = {"ct": "0000", "ct_contrast": "0001"}
 44
 45COHORTS = {"train": ["Internal_Cohort"], "test": [f"Testing_Cohort_{i}" for i in range(1, 5)]}
 46
 47SPLIT_DIRS = {"train": "Tr", "test": "Ts"}
 48
 49
 50def _unzip_with_password(zip_path, dst, password):
 51    try:
 52        with zipfile.ZipFile(zip_path) as f:
 53            f.extractall(dst, pwd=password.encode())
 54    except (NotImplementedError, RuntimeError) as e:
 55        # The python zipfile module does not support AES encryption, so we fall back to the 7z CLI.
 56        if which("7z") is None:
 57            raise RuntimeError(
 58                f"Could not extract '{zip_path}' with the zipfile module ({e}). Please install the '7z' CLI "
 59                "('conda install -c conda-forge p7zip') or extract the archive manually."
 60            )
 61        run(["7z", "x", f"-o{dst}", f"-p{password}", "-y", zip_path], check=True)
 62
 63
 64def get_segrap2025_lnctv_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 65    """Download the SegRap2025 LN CTV dataset.
 66
 67    Args:
 68        path: Filepath to a folder where the data is downloaded for further processing.
 69        download: Whether to download the data if it is not present.
 70
 71    Returns:
 72        Filepath where the data is stored.
 73    """
 74    data_dir = os.path.join(path, "LNCTVSeg-DataSet")
 75    if os.path.exists(data_dir):
 76        return data_dir
 77
 78    os.makedirs(path, exist_ok=True)
 79    zip_path = os.path.join(path, "LNCTVSeg-DataSet.zip")
 80    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
 81    _unzip_with_password(zip_path, path, PASSWORD)
 82    os.remove(zip_path)
 83
 84    return data_dir
 85
 86
 87def get_segrap2025_lnctv_paths(
 88    path: Union[os.PathLike, str],
 89    split: Literal["train", "test"] = "train",
 90    modality: Literal["ct", "ct_contrast"] = "ct",
 91    download: bool = False,
 92) -> Tuple[List[str], List[str]]:
 93    """Get paths to the SegRap2025 LN CTV data.
 94
 95    Args:
 96        path: Filepath to a folder where the data is downloaded for further processing.
 97        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
 98        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
 99        download: Whether to download the data if it is not present.
100
101    Returns:
102        List of filepaths for the image data.
103        List of filepaths for the label data.
104    """
105    if split not in COHORTS:
106        raise ValueError(f"'{split}' is not a valid split. Choose one of {list(COHORTS.keys())}.")
107    if modality not in MODALITIES:
108        raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(MODALITIES.keys())}.")
109
110    data_dir = get_segrap2025_lnctv_data(path, download)
111    suffix = MODALITIES[modality]
112    image_dir, label_dir = f"images{SPLIT_DIRS[split]}", f"labels{SPLIT_DIRS[split]}"
113
114    raw_paths, label_paths = [], []
115    for cohort in COHORTS[split]:
116        cur_raw_paths = natsorted(glob(os.path.join(data_dir, cohort, image_dir, f"*_{suffix}.nii.gz")))
117        cur_label_paths = [
118            os.path.join(data_dir, cohort, label_dir, os.path.basename(p).replace(f"_{suffix}.nii.gz", ".nii.gz"))
119            for p in cur_raw_paths
120        ]
121        keep = [os.path.exists(p) for p in cur_label_paths]
122        raw_paths.extend(p for p, k in zip(cur_raw_paths, keep) if k)
123        label_paths.extend(p for p, k in zip(cur_label_paths, keep) if k)
124
125    assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths)
126
127    return raw_paths, label_paths
128
129
130def get_segrap2025_lnctv_dataset(
131    path: Union[os.PathLike, str],
132    patch_shape: Tuple[int, ...],
133    split: Literal["train", "test"] = "train",
134    modality: Literal["ct", "ct_contrast"] = "ct",
135    resize_inputs: bool = False,
136    download: bool = False,
137    **kwargs
138) -> Dataset:
139    """Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation.
140
141    Args:
142        path: Filepath to a folder where the data is downloaded for further processing.
143        patch_shape: The patch shape to use for training.
144        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
145        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
146        resize_inputs: Whether to resize inputs to the desired patch shape.
147        download: Whether to download the data if it is not present.
148        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
149
150    Returns:
151        The segmentation dataset.
152    """
153    raw_paths, label_paths = get_segrap2025_lnctv_paths(path, split, modality, download)
154
155    if resize_inputs:
156        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
157        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
158            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
159        )
160
161    return torch_em.default_segmentation_dataset(
162        raw_paths=raw_paths,
163        raw_key="data",
164        label_paths=label_paths,
165        label_key="data",
166        patch_shape=patch_shape,
167        is_seg_dataset=True,
168        **kwargs
169    )
170
171
172def get_segrap2025_lnctv_loader(
173    path: Union[os.PathLike, str],
174    batch_size: int,
175    patch_shape: Tuple[int, ...],
176    split: Literal["train", "test"] = "train",
177    modality: Literal["ct", "ct_contrast"] = "ct",
178    resize_inputs: bool = False,
179    download: bool = False,
180    **kwargs
181) -> DataLoader:
182    """Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation.
183
184    Args:
185        path: Filepath to a folder where the data is downloaded for further processing.
186        batch_size: The batch size for training.
187        patch_shape: The patch shape to use for training.
188        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
189        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
190        resize_inputs: Whether to resize inputs to the desired patch shape.
191        download: Whether to download the data if it is not present.
192        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
193
194    Returns:
195        The DataLoader.
196    """
197    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
198    dataset = get_segrap2025_lnctv_dataset(path, patch_shape, split, modality, resize_inputs, download, **ds_kwargs)
199    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://ndownloader.figshare.com/files/48684664'
CHECKSUM = '995215f98844f1a5718a1b963fe5615dc3723d506a889503617fcd99aea052a1'
PASSWORD = 'lnctvseg@uestc'
MODALITIES = {'ct': '0000', 'ct_contrast': '0001'}
COHORTS = {'train': ['Internal_Cohort'], 'test': ['Testing_Cohort_1', 'Testing_Cohort_2', 'Testing_Cohort_3', 'Testing_Cohort_4']}
SPLIT_DIRS = {'train': 'Tr', 'test': 'Ts'}
def get_segrap2025_lnctv_data(path: Union[os.PathLike, str], download: bool = False) -> str:
65def get_segrap2025_lnctv_data(path: Union[os.PathLike, str], download: bool = False) -> str:
66    """Download the SegRap2025 LN CTV dataset.
67
68    Args:
69        path: Filepath to a folder where the data is downloaded for further processing.
70        download: Whether to download the data if it is not present.
71
72    Returns:
73        Filepath where the data is stored.
74    """
75    data_dir = os.path.join(path, "LNCTVSeg-DataSet")
76    if os.path.exists(data_dir):
77        return data_dir
78
79    os.makedirs(path, exist_ok=True)
80    zip_path = os.path.join(path, "LNCTVSeg-DataSet.zip")
81    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
82    _unzip_with_password(zip_path, path, PASSWORD)
83    os.remove(zip_path)
84
85    return data_dir

Download the SegRap2025 LN CTV dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is stored.

def get_segrap2025_lnctv_paths( path: Union[os.PathLike, str], split: Literal['train', 'test'] = 'train', modality: Literal['ct', 'ct_contrast'] = 'ct', download: bool = False) -> Tuple[List[str], List[str]]:
 88def get_segrap2025_lnctv_paths(
 89    path: Union[os.PathLike, str],
 90    split: Literal["train", "test"] = "train",
 91    modality: Literal["ct", "ct_contrast"] = "ct",
 92    download: bool = False,
 93) -> Tuple[List[str], List[str]]:
 94    """Get paths to the SegRap2025 LN CTV data.
 95
 96    Args:
 97        path: Filepath to a folder where the data is downloaded for further processing.
 98        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
 99        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
100        download: Whether to download the data if it is not present.
101
102    Returns:
103        List of filepaths for the image data.
104        List of filepaths for the label data.
105    """
106    if split not in COHORTS:
107        raise ValueError(f"'{split}' is not a valid split. Choose one of {list(COHORTS.keys())}.")
108    if modality not in MODALITIES:
109        raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(MODALITIES.keys())}.")
110
111    data_dir = get_segrap2025_lnctv_data(path, download)
112    suffix = MODALITIES[modality]
113    image_dir, label_dir = f"images{SPLIT_DIRS[split]}", f"labels{SPLIT_DIRS[split]}"
114
115    raw_paths, label_paths = [], []
116    for cohort in COHORTS[split]:
117        cur_raw_paths = natsorted(glob(os.path.join(data_dir, cohort, image_dir, f"*_{suffix}.nii.gz")))
118        cur_label_paths = [
119            os.path.join(data_dir, cohort, label_dir, os.path.basename(p).replace(f"_{suffix}.nii.gz", ".nii.gz"))
120            for p in cur_raw_paths
121        ]
122        keep = [os.path.exists(p) for p in cur_label_paths]
123        raw_paths.extend(p for p, k in zip(cur_raw_paths, keep) if k)
124        label_paths.extend(p for p, k in zip(cur_label_paths, keep) if k)
125
126    assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths)
127
128    return raw_paths, label_paths

Get paths to the SegRap2025 LN CTV data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_segrap2025_lnctv_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], split: Literal['train', 'test'] = 'train', modality: Literal['ct', 'ct_contrast'] = 'ct', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
131def get_segrap2025_lnctv_dataset(
132    path: Union[os.PathLike, str],
133    patch_shape: Tuple[int, ...],
134    split: Literal["train", "test"] = "train",
135    modality: Literal["ct", "ct_contrast"] = "ct",
136    resize_inputs: bool = False,
137    download: bool = False,
138    **kwargs
139) -> Dataset:
140    """Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation.
141
142    Args:
143        path: Filepath to a folder where the data is downloaded for further processing.
144        patch_shape: The patch shape to use for training.
145        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
146        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
147        resize_inputs: Whether to resize inputs to the desired patch shape.
148        download: Whether to download the data if it is not present.
149        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
150
151    Returns:
152        The segmentation dataset.
153    """
154    raw_paths, label_paths = get_segrap2025_lnctv_paths(path, split, modality, download)
155
156    if resize_inputs:
157        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
158        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
159            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
160        )
161
162    return torch_em.default_segmentation_dataset(
163        raw_paths=raw_paths,
164        raw_key="data",
165        label_paths=label_paths,
166        label_key="data",
167        patch_shape=patch_shape,
168        is_seg_dataset=True,
169        **kwargs
170    )

Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_segrap2025_lnctv_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], split: Literal['train', 'test'] = 'train', modality: Literal['ct', 'ct_contrast'] = 'ct', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
173def get_segrap2025_lnctv_loader(
174    path: Union[os.PathLike, str],
175    batch_size: int,
176    patch_shape: Tuple[int, ...],
177    split: Literal["train", "test"] = "train",
178    modality: Literal["ct", "ct_contrast"] = "ct",
179    resize_inputs: bool = False,
180    download: bool = False,
181    **kwargs
182) -> DataLoader:
183    """Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation.
184
185    Args:
186        path: Filepath to a folder where the data is downloaded for further processing.
187        batch_size: The batch size for training.
188        patch_shape: The patch shape to use for training.
189        split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
190        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
191        resize_inputs: Whether to resize inputs to the desired patch shape.
192        download: Whether to download the data if it is not present.
193        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
194
195    Returns:
196        The DataLoader.
197    """
198    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
199    dataset = get_segrap2025_lnctv_dataset(path, patch_shape, split, modality, resize_inputs, download, **ds_kwargs)
200    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.