torch_em.data.datasets.medical.segrap2025
The SegRap2025 dataset contains annotations for lymph node clinical target volume (LN CTV) segmentation in head and neck CT scans of nasopharyngeal carcinoma patients.
It is the "Task02: LN CTV Segmentation" part of the SegRap2025 challenge (https://hilab-git.github.io/SegRap2025_Challenge) and is unrelated to the organs-at-risk data of SegRap2023 covered in 'segrap.py': it comprises a different set of 440 CT scans from 262 patients, collected across 5 cohorts (4 centers), each with a pre-aligned pair of a non-contrast and a contrast-enhanced CT scan. 120 patients from the internal cohort provide the official training split (labelled), the remaining 4 testing cohorts (142 patients) are distributed with labels as part of this public release too, and are exposed here as the 'test' split.
NOTE: The label legend is not documented in the released data. Verified on the data: the label volumes contain the ids 0 to 6, i.e. background and 6 disjoint LN CTV sub-levels.
The dataset is a redistribution of the official SegRap2025 Task02 training data (https://doi.org/10.6084/m9.figshare.26793622, CC BY 4.0) via Figshare. The archive is password protected, the password ('lnctvseg@uestc') is published on the challenge dataset page.
This dataset is from the publication https://doi.org/10.48550/arXiv.2601.20575. Please cite it if you use this dataset in your research.
1"""The SegRap2025 dataset contains annotations for lymph node clinical target volume (LN CTV) segmentation 2in head and neck CT scans of nasopharyngeal carcinoma patients. 3 4It is the "Task02: LN CTV Segmentation" part of the SegRap2025 challenge 5(https://hilab-git.github.io/SegRap2025_Challenge) and is unrelated to the organs-at-risk data of SegRap2023 6covered in 'segrap.py': it comprises a different set of 7440 CT scans from 262 patients, collected across 5 cohorts (4 centers), each with a pre-aligned pair of a 8non-contrast and a contrast-enhanced CT scan. 120 patients from the internal cohort provide the official training 9split (labelled), the remaining 4 testing cohorts (142 patients) are distributed with labels as part of this 10public release too, and are exposed here as the 'test' split. 11 12NOTE: The label legend is not documented in the released data. Verified on the data: the label volumes contain 13the ids 0 to 6, i.e. background and 6 disjoint LN CTV sub-levels. 14 15The dataset is a redistribution of the official SegRap2025 Task02 training data 16(https://doi.org/10.6084/m9.figshare.26793622, CC BY 4.0) via Figshare. The archive is password protected, the 17password ('lnctvseg@uestc') is published on the challenge dataset page. 18 19This dataset is from the publication https://doi.org/10.48550/arXiv.2601.20575. 20Please cite it if you use this dataset in your research. 21""" 22 23import os 24import zipfile 25from glob import glob 26from shutil import which 27from subprocess import run 28from natsort import natsorted 29from typing import Union, Tuple, Literal, List 30 31from torch.utils.data import Dataset, DataLoader 32 33import torch_em 34 35from .. import util 36 37 38URL = "https://ndownloader.figshare.com/files/48684664" 39CHECKSUM = "995215f98844f1a5718a1b963fe5615dc3723d506a889503617fcd99aea052a1" 40 41PASSWORD = "lnctvseg@uestc" 42 43MODALITIES = {"ct": "0000", "ct_contrast": "0001"} 44 45COHORTS = {"train": ["Internal_Cohort"], "test": [f"Testing_Cohort_{i}" for i in range(1, 5)]} 46 47SPLIT_DIRS = {"train": "Tr", "test": "Ts"} 48 49 50def _unzip_with_password(zip_path, dst, password): 51 try: 52 with zipfile.ZipFile(zip_path) as f: 53 f.extractall(dst, pwd=password.encode()) 54 except (NotImplementedError, RuntimeError) as e: 55 # The python zipfile module does not support AES encryption, so we fall back to the 7z CLI. 56 if which("7z") is None: 57 raise RuntimeError( 58 f"Could not extract '{zip_path}' with the zipfile module ({e}). Please install the '7z' CLI " 59 "('conda install -c conda-forge p7zip') or extract the archive manually." 60 ) 61 run(["7z", "x", f"-o{dst}", f"-p{password}", "-y", zip_path], check=True) 62 63 64def get_segrap2025_lnctv_data(path: Union[os.PathLike, str], download: bool = False) -> str: 65 """Download the SegRap2025 LN CTV dataset. 66 67 Args: 68 path: Filepath to a folder where the data is downloaded for further processing. 69 download: Whether to download the data if it is not present. 70 71 Returns: 72 Filepath where the data is stored. 73 """ 74 data_dir = os.path.join(path, "LNCTVSeg-DataSet") 75 if os.path.exists(data_dir): 76 return data_dir 77 78 os.makedirs(path, exist_ok=True) 79 zip_path = os.path.join(path, "LNCTVSeg-DataSet.zip") 80 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 81 _unzip_with_password(zip_path, path, PASSWORD) 82 os.remove(zip_path) 83 84 return data_dir 85 86 87def get_segrap2025_lnctv_paths( 88 path: Union[os.PathLike, str], 89 split: Literal["train", "test"] = "train", 90 modality: Literal["ct", "ct_contrast"] = "ct", 91 download: bool = False, 92) -> Tuple[List[str], List[str]]: 93 """Get paths to the SegRap2025 LN CTV data. 94 95 Args: 96 path: Filepath to a folder where the data is downloaded for further processing. 97 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 98 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 99 download: Whether to download the data if it is not present. 100 101 Returns: 102 List of filepaths for the image data. 103 List of filepaths for the label data. 104 """ 105 if split not in COHORTS: 106 raise ValueError(f"'{split}' is not a valid split. Choose one of {list(COHORTS.keys())}.") 107 if modality not in MODALITIES: 108 raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(MODALITIES.keys())}.") 109 110 data_dir = get_segrap2025_lnctv_data(path, download) 111 suffix = MODALITIES[modality] 112 image_dir, label_dir = f"images{SPLIT_DIRS[split]}", f"labels{SPLIT_DIRS[split]}" 113 114 raw_paths, label_paths = [], [] 115 for cohort in COHORTS[split]: 116 cur_raw_paths = natsorted(glob(os.path.join(data_dir, cohort, image_dir, f"*_{suffix}.nii.gz"))) 117 cur_label_paths = [ 118 os.path.join(data_dir, cohort, label_dir, os.path.basename(p).replace(f"_{suffix}.nii.gz", ".nii.gz")) 119 for p in cur_raw_paths 120 ] 121 keep = [os.path.exists(p) for p in cur_label_paths] 122 raw_paths.extend(p for p, k in zip(cur_raw_paths, keep) if k) 123 label_paths.extend(p for p, k in zip(cur_label_paths, keep) if k) 124 125 assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths) 126 127 return raw_paths, label_paths 128 129 130def get_segrap2025_lnctv_dataset( 131 path: Union[os.PathLike, str], 132 patch_shape: Tuple[int, ...], 133 split: Literal["train", "test"] = "train", 134 modality: Literal["ct", "ct_contrast"] = "ct", 135 resize_inputs: bool = False, 136 download: bool = False, 137 **kwargs 138) -> Dataset: 139 """Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation. 140 141 Args: 142 path: Filepath to a folder where the data is downloaded for further processing. 143 patch_shape: The patch shape to use for training. 144 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 145 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 146 resize_inputs: Whether to resize inputs to the desired patch shape. 147 download: Whether to download the data if it is not present. 148 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 149 150 Returns: 151 The segmentation dataset. 152 """ 153 raw_paths, label_paths = get_segrap2025_lnctv_paths(path, split, modality, download) 154 155 if resize_inputs: 156 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 157 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 158 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 159 ) 160 161 return torch_em.default_segmentation_dataset( 162 raw_paths=raw_paths, 163 raw_key="data", 164 label_paths=label_paths, 165 label_key="data", 166 patch_shape=patch_shape, 167 is_seg_dataset=True, 168 **kwargs 169 ) 170 171 172def get_segrap2025_lnctv_loader( 173 path: Union[os.PathLike, str], 174 batch_size: int, 175 patch_shape: Tuple[int, ...], 176 split: Literal["train", "test"] = "train", 177 modality: Literal["ct", "ct_contrast"] = "ct", 178 resize_inputs: bool = False, 179 download: bool = False, 180 **kwargs 181) -> DataLoader: 182 """Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation. 183 184 Args: 185 path: Filepath to a folder where the data is downloaded for further processing. 186 batch_size: The batch size for training. 187 patch_shape: The patch shape to use for training. 188 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 189 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 190 resize_inputs: Whether to resize inputs to the desired patch shape. 191 download: Whether to download the data if it is not present. 192 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 193 194 Returns: 195 The DataLoader. 196 """ 197 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 198 dataset = get_segrap2025_lnctv_dataset(path, patch_shape, split, modality, resize_inputs, download, **ds_kwargs) 199 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
65def get_segrap2025_lnctv_data(path: Union[os.PathLike, str], download: bool = False) -> str: 66 """Download the SegRap2025 LN CTV dataset. 67 68 Args: 69 path: Filepath to a folder where the data is downloaded for further processing. 70 download: Whether to download the data if it is not present. 71 72 Returns: 73 Filepath where the data is stored. 74 """ 75 data_dir = os.path.join(path, "LNCTVSeg-DataSet") 76 if os.path.exists(data_dir): 77 return data_dir 78 79 os.makedirs(path, exist_ok=True) 80 zip_path = os.path.join(path, "LNCTVSeg-DataSet.zip") 81 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 82 _unzip_with_password(zip_path, path, PASSWORD) 83 os.remove(zip_path) 84 85 return data_dir
Download the SegRap2025 LN CTV dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is stored.
88def get_segrap2025_lnctv_paths( 89 path: Union[os.PathLike, str], 90 split: Literal["train", "test"] = "train", 91 modality: Literal["ct", "ct_contrast"] = "ct", 92 download: bool = False, 93) -> Tuple[List[str], List[str]]: 94 """Get paths to the SegRap2025 LN CTV data. 95 96 Args: 97 path: Filepath to a folder where the data is downloaded for further processing. 98 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 99 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 100 download: Whether to download the data if it is not present. 101 102 Returns: 103 List of filepaths for the image data. 104 List of filepaths for the label data. 105 """ 106 if split not in COHORTS: 107 raise ValueError(f"'{split}' is not a valid split. Choose one of {list(COHORTS.keys())}.") 108 if modality not in MODALITIES: 109 raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(MODALITIES.keys())}.") 110 111 data_dir = get_segrap2025_lnctv_data(path, download) 112 suffix = MODALITIES[modality] 113 image_dir, label_dir = f"images{SPLIT_DIRS[split]}", f"labels{SPLIT_DIRS[split]}" 114 115 raw_paths, label_paths = [], [] 116 for cohort in COHORTS[split]: 117 cur_raw_paths = natsorted(glob(os.path.join(data_dir, cohort, image_dir, f"*_{suffix}.nii.gz"))) 118 cur_label_paths = [ 119 os.path.join(data_dir, cohort, label_dir, os.path.basename(p).replace(f"_{suffix}.nii.gz", ".nii.gz")) 120 for p in cur_raw_paths 121 ] 122 keep = [os.path.exists(p) for p in cur_label_paths] 123 raw_paths.extend(p for p, k in zip(cur_raw_paths, keep) if k) 124 label_paths.extend(p for p, k in zip(cur_label_paths, keep) if k) 125 126 assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths) 127 128 return raw_paths, label_paths
Get paths to the SegRap2025 LN CTV data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
- modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
131def get_segrap2025_lnctv_dataset( 132 path: Union[os.PathLike, str], 133 patch_shape: Tuple[int, ...], 134 split: Literal["train", "test"] = "train", 135 modality: Literal["ct", "ct_contrast"] = "ct", 136 resize_inputs: bool = False, 137 download: bool = False, 138 **kwargs 139) -> Dataset: 140 """Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation. 141 142 Args: 143 path: Filepath to a folder where the data is downloaded for further processing. 144 patch_shape: The patch shape to use for training. 145 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 146 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 147 resize_inputs: Whether to resize inputs to the desired patch shape. 148 download: Whether to download the data if it is not present. 149 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 150 151 Returns: 152 The segmentation dataset. 153 """ 154 raw_paths, label_paths = get_segrap2025_lnctv_paths(path, split, modality, download) 155 156 if resize_inputs: 157 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 158 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 159 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 160 ) 161 162 return torch_em.default_segmentation_dataset( 163 raw_paths=raw_paths, 164 raw_key="data", 165 label_paths=label_paths, 166 label_key="data", 167 patch_shape=patch_shape, 168 is_seg_dataset=True, 169 **kwargs 170 )
Get the SegRap2025 LN CTV dataset for lymph node clinical target volume segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
- modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
173def get_segrap2025_lnctv_loader( 174 path: Union[os.PathLike, str], 175 batch_size: int, 176 patch_shape: Tuple[int, ...], 177 split: Literal["train", "test"] = "train", 178 modality: Literal["ct", "ct_contrast"] = "ct", 179 resize_inputs: bool = False, 180 download: bool = False, 181 **kwargs 182) -> DataLoader: 183 """Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation. 184 185 Args: 186 path: Filepath to a folder where the data is downloaded for further processing. 187 batch_size: The batch size for training. 188 patch_shape: The patch shape to use for training. 189 split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts). 190 modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced). 191 resize_inputs: Whether to resize inputs to the desired patch shape. 192 download: Whether to download the data if it is not present. 193 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 194 195 Returns: 196 The DataLoader. 197 """ 198 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 199 dataset = get_segrap2025_lnctv_dataset(path, patch_shape, split, modality, resize_inputs, download, **ds_kwargs) 200 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the SegRap2025 LN CTV dataloader for lymph node clinical target volume segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- split: The choice of data split. Either 'train' (the internal cohort) or 'test' (the 4 external cohorts).
- modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.