torch_em.data.datasets.medical.segrap

The SegRap dataset contains annotations for organ-at-risk (OAR) and gross tumor volume (GTV) segmentation in head and neck CT scans of nasopharyngeal carcinoma patients.

It comprises the training set of the SegRap2023 challenge (https://segrap2023.grand-challenge.org): 120 patients with a pre-aligned pair of a non-contrast and a contrast-enhanced CT scan, and annotations for two tasks, selected with the 'task' argument: 'oars' (Task001) are the 45 OARs, 'gtv' (Task002) are the primary gross tumor volume (GTVp) and the involved metastatic lymph nodes (GTVnd), see GTV_LABEL_IDS. The ids were verified on the data: the Task002 label volumes contain the ids 0 (background), 1 (GTVp) and 2 (GTVnd).

NOTE: The label legend is as follows. Since some of the 45 OARs are nested (e.g. the hippocampi inside the temporal lobes, the cochleae inside the middle ears), the challenge distributes the annotations as a single label volume with 54 disjoint sub-parts, where each overlap of two OARs gets its own id. The ids are (see SUBPART_NAMES and LABEL_IDS):

  • 1: Brain, 2: BrainStem, 3: Chiasm, 4/5: TemporalLobe_L/R, 6/7: TemporalLobe_Hippocampus_L/R, 8/9: Hippocampus_L/R, 10/11: Eye_L/R, 12/13: Lens_L/R, 14/15: OpticNerve_L/R, 16/17: MiddleEar_L/R, 18/19: IAC_L/R, 20/21: MiddleEar_TympanicCavity_L/R, 22/23: TympanicCavity_L/R, 24/25: MiddleEar_VestibulSemi_L/R, 26/27: VestibulSemi_L/R, 28/29: Cochlea_L/R, 30/31: MiddleEar_ETbone_L/R, 32/33: ETbone_L/R, 34: Pituitary, 35: OralCavity, 36/37: Mandible_L/R, 38/39: Submandibular_L/R, 40/41: Parotid_L/R, 42/43: Mastoid_L/R, 44/45: TMjoint_L/R, 46: SpinalCord, 47: Esophagus, 48: Larynx, 49: Larynx_Glottic, 50: Larynx_Supraglot, 51: Larynx_PharynxConst, 52: PharynxConst, 53: Thyroid, 54: Trachea OAR_TO_LABEL_IDS maps each of the 45 OARs to the ids it consists of. It is taken from the official post-processing code of the challenge (https://github.com/HiLab-git/SegRap2023/blob/main/Tutorial/postprocessing.py). The ids were verified on the data: the label volumes contain the ids 0 to 54.

The data is a redistribution of the official challenge data at https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite (CC BY-NC 4.0), which holds the unchanged images and Task001 labels of all 120 training cases, renamed after the case ids. The official release at https://segrap2023.grand-challenge.org/dataset/ requires a signed end user agreement, so please make sure that you are allowed to use the data for your purpose.

NOTE: The official release can also be used. Download 'SegRap2023_Training_Set_120cases.zip' and 'SegRap2023_Training_Set_120cases_OneHot_Labels.zip' as described on the dataset page and extract them into '', such that '/SegRap2023_Training_Set_120cases/segrap_XXXX/image.nii.gz' (and 'image_contrast.nii.gz') and '/SegRap2023_Training_Set_120cases_OneHot_Labels/Task001/segrap_XXXX.nii.gz' exist. This dataset will then use the official data instead of the redistribution.

This dataset is from the publication https://doi.org/10.1016/j.media.2024.103447. Please cite it if you use this dataset in your research.

  1"""The SegRap dataset contains annotations for organ-at-risk (OAR) and gross tumor volume (GTV)
  2segmentation in head and neck CT scans of nasopharyngeal carcinoma patients.
  3
  4It comprises the training set of the SegRap2023 challenge (https://segrap2023.grand-challenge.org): 120 patients
  5with a pre-aligned pair of a non-contrast and a contrast-enhanced CT scan, and annotations for two tasks,
  6selected with the 'task' argument: 'oars' (Task001) are the 45 OARs, 'gtv' (Task002) are the primary gross
  7tumor volume (GTVp) and the involved metastatic lymph nodes (GTVnd), see `GTV_LABEL_IDS`. The ids were
  8verified on the data: the Task002 label volumes contain the ids 0 (background), 1 (GTVp) and 2 (GTVnd).
  9
 10NOTE: The label legend is as follows. Since some of the 45 OARs are nested (e.g. the hippocampi inside the
 11temporal lobes, the cochleae inside the middle ears), the challenge distributes the annotations as a single
 12label volume with 54 disjoint sub-parts, where each overlap of two OARs gets its own id. The ids are
 13(see `SUBPART_NAMES` and `LABEL_IDS`):
 14- 1: Brain, 2: BrainStem, 3: Chiasm, 4/5: TemporalLobe_L/R, 6/7: TemporalLobe_Hippocampus_L/R,
 15  8/9: Hippocampus_L/R, 10/11: Eye_L/R, 12/13: Lens_L/R, 14/15: OpticNerve_L/R, 16/17: MiddleEar_L/R,
 16  18/19: IAC_L/R, 20/21: MiddleEar_TympanicCavity_L/R, 22/23: TympanicCavity_L/R,
 17  24/25: MiddleEar_VestibulSemi_L/R, 26/27: VestibulSemi_L/R, 28/29: Cochlea_L/R,
 18  30/31: MiddleEar_ETbone_L/R, 32/33: ETbone_L/R, 34: Pituitary, 35: OralCavity, 36/37: Mandible_L/R,
 19  38/39: Submandibular_L/R, 40/41: Parotid_L/R, 42/43: Mastoid_L/R, 44/45: TMjoint_L/R, 46: SpinalCord,
 20  47: Esophagus, 48: Larynx, 49: Larynx_Glottic, 50: Larynx_Supraglot, 51: Larynx_PharynxConst,
 21  52: PharynxConst, 53: Thyroid, 54: Trachea
 22`OAR_TO_LABEL_IDS` maps each of the 45 OARs to the ids it consists of. It is taken from the official
 23post-processing code of the challenge (https://github.com/HiLab-git/SegRap2023/blob/main/Tutorial/postprocessing.py).
 24The ids were verified on the data: the label volumes contain the ids 0 to 54.
 25
 26The data is a redistribution of the official challenge data at
 27https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite (CC BY-NC 4.0), which holds the unchanged images
 28and Task001 labels of all 120 training cases, renamed after the case ids. The official release at
 29https://segrap2023.grand-challenge.org/dataset/ requires a signed end user agreement, so please make sure
 30that you are allowed to use the data for your purpose.
 31
 32NOTE: The official release can also be used. Download 'SegRap2023_Training_Set_120cases.zip' and
 33'SegRap2023_Training_Set_120cases_OneHot_Labels.zip' as described on the dataset page and extract them into
 34'<path>', such that '<path>/SegRap2023_Training_Set_120cases/segrap_XXXX/image.nii.gz' (and
 35'image_contrast.nii.gz') and '<path>/SegRap2023_Training_Set_120cases_OneHot_Labels/Task001/segrap_XXXX.nii.gz'
 36exist. This dataset will then use the official data instead of the redistribution.
 37
 38This dataset is from the publication https://doi.org/10.1016/j.media.2024.103447.
 39Please cite it if you use this dataset in your research.
 40"""
 41
 42import os
 43from glob import glob
 44from natsort import natsorted
 45from typing import Union, Tuple, Literal, List
 46
 47from torch.utils.data import Dataset, DataLoader
 48
 49import torch_em
 50
 51from .. import util
 52
 53
 54URLS = {
 55    "ct": "https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Images-CT.zip",
 56    "ct_contrast": "https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Images-contrastCT.zip",
 57    "labels_oars": "https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Masks-Task1.zip",
 58    "labels_gtv": "https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Masks-Task2.zip",
 59}
 60
 61CHECKSUMS = {
 62    "ct": "1e7b849c5f0296200e5ad9503eeeb9b2a8b3360297078d0ce016c8f8335534ef",
 63    "ct_contrast": "000cc8bf9041c49f7e28ad3e183f7af3f9e701b1842031fb17e21a89db0a8f6a",
 64    "labels_oars": "3441d2bd56ecff4485b99a197b36251b09d2030468840f9399419b517691d297",
 65    "labels_gtv": "d412baf209f7a2d2432f05d1f39b5ce1944e1ff2d2b6ccfc4f727626a02f3d55",
 66}
 67
 68FOLDER_NAMES = {
 69    "ct": "Images-CT",
 70    "ct_contrast": "Images-contrastCT",
 71    "labels_oars": "Masks-Task1",
 72    "labels_gtv": "Masks-Task2",
 73}
 74
 75# The label component and the label folder of the official release, per task.
 76TASKS = {"oars": ("labels_oars", "Task001"), "gtv": ("labels_gtv", "Task002")}
 77
 78GTV_LABEL_IDS = {"background": 0, "GTVp": 1, "GTVnd": 2}
 79
 80# The file names of the two scans in the official release.
 81OFFICIAL_FILE_NAMES = {"ct": "image.nii.gz", "ct_contrast": "image_contrast.nii.gz"}
 82
 83SUBPART_NAMES = [
 84    "Brain", "BrainStem", "Chiasm", "TemporalLobe_L", "TemporalLobe_R", "TemporalLobe_Hippocampus_L",
 85    "TemporalLobe_Hippocampus_R", "Hippocampus_L", "Hippocampus_R", "Eye_L", "Eye_R", "Lens_L", "Lens_R",
 86    "OpticNerve_L", "OpticNerve_R", "MiddleEar_L", "MiddleEar_R", "IAC_L", "IAC_R",
 87    "MiddleEar_TympanicCavity_L", "MiddleEar_TympanicCavity_R", "TympanicCavity_L", "TympanicCavity_R",
 88    "MiddleEar_VestibulSemi_L", "MiddleEar_VestibulSemi_R", "VestibulSemi_L", "VestibulSemi_R", "Cochlea_L",
 89    "Cochlea_R", "MiddleEar_ETbone_L", "MiddleEar_ETbone_R", "ETbone_L", "ETbone_R", "Pituitary", "OralCavity",
 90    "Mandible_L", "Mandible_R", "Submandibular_L", "Submandibular_R", "Parotid_L", "Parotid_R", "Mastoid_L",
 91    "Mastoid_R", "TMjoint_L", "TMjoint_R", "SpinalCord", "Esophagus", "Larynx", "Larynx_Glottic",
 92    "Larynx_Supraglot", "Larynx_PharynxConst", "PharynxConst", "Thyroid", "Trachea",
 93]
 94
 95LABEL_IDS = {"background": 0, **{name: i + 1 for i, name in enumerate(SUBPART_NAMES)}}
 96
 97OAR_TO_LABEL_IDS = {
 98    "Brain": [1, 2, 3, 4, 5, 6, 7, 8, 9],
 99    "BrainStem": [2],
100    "Chiasm": [3],
101    "TemporalLobe_L": [4, 6],
102    "TemporalLobe_R": [5, 7],
103    "Hippocampus_L": [8, 6],
104    "Hippocampus_R": [9, 7],
105    "Eye_L": [10, 12],
106    "Eye_R": [11, 13],
107    "Lens_L": [12],
108    "Lens_R": [13],
109    "OpticNerve_L": [14],
110    "OpticNerve_R": [15],
111    "MiddleEar_L": [18, 16, 20, 24, 28, 30],
112    "MiddleEar_R": [19, 17, 21, 25, 29, 31],
113    "IAC_L": [18],
114    "IAC_R": [19],
115    "TympanicCavity_L": [22, 20],
116    "TympanicCavity_R": [23, 21],
117    "VestibulSemi_L": [26, 24],
118    "VestibulSemi_R": [27, 25],
119    "Cochlea_L": [28],
120    "Cochlea_R": [29],
121    "ETbone_L": [32, 30],
122    "ETbone_R": [33, 31],
123    "Pituitary": [34],
124    "OralCavity": [35],
125    "Mandible_L": [36],
126    "Mandible_R": [37],
127    "Submandibular_L": [38],
128    "Submandibular_R": [39],
129    "Parotid_L": [40],
130    "Parotid_R": [41],
131    "Mastoid_L": [42],
132    "Mastoid_R": [43],
133    "TMjoint_L": [44],
134    "TMjoint_R": [45],
135    "SpinalCord": [46],
136    "Esophagus": [47],
137    "Larynx": [48, 49, 50, 51],
138    "Larynx_Glottic": [49],
139    "Larynx_Supraglot": [50],
140    "PharynxConst": [51, 52],
141    "Thyroid": [53],
142    "Trachea": [54],
143}
144
145
146def _get_official_case_dirs(path):
147    case_dirs = glob(os.path.join(path, "**", "SegRap2023_Training_Set_120cases", "segrap_*"), recursive=True)
148    return natsorted([p for p in case_dirs if os.path.isdir(p)])
149
150
151def _download_component(path, name, download):
152    data_dir = os.path.join(path, FOLDER_NAMES[name])
153    if os.path.exists(data_dir):
154        return data_dir
155
156    os.makedirs(path, exist_ok=True)
157    zip_path = os.path.join(path, f"{FOLDER_NAMES[name]}.zip")
158    util.download_source(path=zip_path, url=URLS[name], download=download, checksum=CHECKSUMS[name])
159    util.unzip(zip_path=zip_path, dst=path)
160
161    return data_dir
162
163
164def get_segrap_data(
165    path: Union[os.PathLike, str],
166    modality: Literal["ct", "ct_contrast"] = "ct",
167    download: bool = False,
168    task: Literal["oars", "gtv"] = "oars",
169) -> str:
170    """Download the SegRap dataset.
171
172    Args:
173        path: Filepath to a folder where the data is downloaded for further processing.
174        modality: The CT scan to download. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
175        download: Whether to download the data if it is not present.
176        task: The annotations to download. Either 'oars' (Task001) or 'gtv' (Task002).
177
178    Returns:
179        Filepath where the data is stored.
180    """
181    if modality not in OFFICIAL_FILE_NAMES:
182        raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(OFFICIAL_FILE_NAMES)}.")
183    if task not in TASKS:
184        raise ValueError(f"'{task}' is not a valid task. Choose one of {list(TASKS)}.")
185
186    if len(_get_official_case_dirs(path)) > 0:  # The official data was downloaded manually.
187        return path
188
189    _download_component(path, TASKS[task][0], download)
190    _download_component(path, modality, download)
191
192    return path
193
194
195def get_segrap_paths(
196    path: Union[os.PathLike, str],
197    modality: Literal["ct", "ct_contrast"] = "ct",
198    download: bool = False,
199    task: Literal["oars", "gtv"] = "oars",
200) -> Tuple[List[str], List[str]]:
201    """Get paths to the SegRap data.
202
203    Args:
204        path: Filepath to a folder where the data is downloaded for further processing.
205        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
206        download: Whether to download the data if it is not present.
207        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
208
209    Returns:
210        List of filepaths for the image data.
211        List of filepaths for the label data.
212    """
213    data_dir = get_segrap_data(path, modality, download, task)
214    label_component, official_label_dir = TASKS[task]
215
216    case_dirs = _get_official_case_dirs(data_dir)
217    if len(case_dirs) > 0:  # The official layout, with one folder per case and the labels in a separate folder.
218        label_dir = os.path.join(os.path.split(os.path.split(case_dirs[0])[0])[0], official_label_dir)
219        raw_paths = [os.path.join(p, OFFICIAL_FILE_NAMES[modality]) for p in case_dirs]
220        label_paths = [os.path.join(label_dir, f"{os.path.basename(p)}.nii.gz") for p in case_dirs]
221    else:  # The redistributed layout, with the files named after the case ids.
222        raw_paths = natsorted(glob(os.path.join(data_dir, FOLDER_NAMES[modality], "*.nii.gz")))
223        label_paths = [
224            os.path.join(data_dir, FOLDER_NAMES[label_component], os.path.basename(p)) for p in raw_paths
225        ]
226
227    assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths + label_paths)
228
229    return raw_paths, label_paths
230
231
232def get_segrap_dataset(
233    path: Union[os.PathLike, str],
234    patch_shape: Tuple[int, ...],
235    modality: Literal["ct", "ct_contrast"] = "ct",
236    resize_inputs: bool = False,
237    download: bool = False,
238    task: Literal["oars", "gtv"] = "oars",
239    **kwargs
240) -> Dataset:
241    """Get the SegRap dataset for organ-at-risk or GTV segmentation.
242
243    Args:
244        path: Filepath to a folder where the data is downloaded for further processing.
245        patch_shape: The patch shape to use for training.
246        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
247        resize_inputs: Whether to resize inputs to the desired patch shape.
248        download: Whether to download the data if it is not present.
249        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
250        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
251
252    Returns:
253        The segmentation dataset.
254    """
255    raw_paths, label_paths = get_segrap_paths(path, modality, download, task)
256
257    if resize_inputs:
258        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
259        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
260            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
261        )
262
263    return torch_em.default_segmentation_dataset(
264        raw_paths=raw_paths,
265        raw_key="data",
266        label_paths=label_paths,
267        label_key="data",
268        patch_shape=patch_shape,
269        is_seg_dataset=True,
270        **kwargs
271    )
272
273
274def get_segrap_loader(
275    path: Union[os.PathLike, str],
276    batch_size: int,
277    patch_shape: Tuple[int, ...],
278    modality: Literal["ct", "ct_contrast"] = "ct",
279    resize_inputs: bool = False,
280    download: bool = False,
281    task: Literal["oars", "gtv"] = "oars",
282    **kwargs
283) -> DataLoader:
284    """Get the SegRap dataloader for organ-at-risk or GTV segmentation.
285
286    Args:
287        path: Filepath to a folder where the data is downloaded for further processing.
288        batch_size: The batch size for training.
289        patch_shape: The patch shape to use for training.
290        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
291        resize_inputs: Whether to resize inputs to the desired patch shape.
292        download: Whether to download the data if it is not present.
293        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
294        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
295
296    Returns:
297        The DataLoader.
298    """
299    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
300    dataset = get_segrap_dataset(path, patch_shape, modality, resize_inputs, download, task, **ds_kwargs)
301    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URLS = {'ct': 'https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Images-CT.zip', 'ct_contrast': 'https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Images-contrastCT.zip', 'labels_oars': 'https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Masks-Task1.zip', 'labels_gtv': 'https://huggingface.co/datasets/YongchengYAO/SegRap23-Lite/resolve/main/Masks-Task2.zip'}
CHECKSUMS = {'ct': '1e7b849c5f0296200e5ad9503eeeb9b2a8b3360297078d0ce016c8f8335534ef', 'ct_contrast': '000cc8bf9041c49f7e28ad3e183f7af3f9e701b1842031fb17e21a89db0a8f6a', 'labels_oars': '3441d2bd56ecff4485b99a197b36251b09d2030468840f9399419b517691d297', 'labels_gtv': 'd412baf209f7a2d2432f05d1f39b5ce1944e1ff2d2b6ccfc4f727626a02f3d55'}
FOLDER_NAMES = {'ct': 'Images-CT', 'ct_contrast': 'Images-contrastCT', 'labels_oars': 'Masks-Task1', 'labels_gtv': 'Masks-Task2'}
TASKS = {'oars': ('labels_oars', 'Task001'), 'gtv': ('labels_gtv', 'Task002')}
GTV_LABEL_IDS = {'background': 0, 'GTVp': 1, 'GTVnd': 2}
OFFICIAL_FILE_NAMES = {'ct': 'image.nii.gz', 'ct_contrast': 'image_contrast.nii.gz'}
SUBPART_NAMES = ['Brain', 'BrainStem', 'Chiasm', 'TemporalLobe_L', 'TemporalLobe_R', 'TemporalLobe_Hippocampus_L', 'TemporalLobe_Hippocampus_R', 'Hippocampus_L', 'Hippocampus_R', 'Eye_L', 'Eye_R', 'Lens_L', 'Lens_R', 'OpticNerve_L', 'OpticNerve_R', 'MiddleEar_L', 'MiddleEar_R', 'IAC_L', 'IAC_R', 'MiddleEar_TympanicCavity_L', 'MiddleEar_TympanicCavity_R', 'TympanicCavity_L', 'TympanicCavity_R', 'MiddleEar_VestibulSemi_L', 'MiddleEar_VestibulSemi_R', 'VestibulSemi_L', 'VestibulSemi_R', 'Cochlea_L', 'Cochlea_R', 'MiddleEar_ETbone_L', 'MiddleEar_ETbone_R', 'ETbone_L', 'ETbone_R', 'Pituitary', 'OralCavity', 'Mandible_L', 'Mandible_R', 'Submandibular_L', 'Submandibular_R', 'Parotid_L', 'Parotid_R', 'Mastoid_L', 'Mastoid_R', 'TMjoint_L', 'TMjoint_R', 'SpinalCord', 'Esophagus', 'Larynx', 'Larynx_Glottic', 'Larynx_Supraglot', 'Larynx_PharynxConst', 'PharynxConst', 'Thyroid', 'Trachea']
LABEL_IDS = {'background': 0, 'Brain': 1, 'BrainStem': 2, 'Chiasm': 3, 'TemporalLobe_L': 4, 'TemporalLobe_R': 5, 'TemporalLobe_Hippocampus_L': 6, 'TemporalLobe_Hippocampus_R': 7, 'Hippocampus_L': 8, 'Hippocampus_R': 9, 'Eye_L': 10, 'Eye_R': 11, 'Lens_L': 12, 'Lens_R': 13, 'OpticNerve_L': 14, 'OpticNerve_R': 15, 'MiddleEar_L': 16, 'MiddleEar_R': 17, 'IAC_L': 18, 'IAC_R': 19, 'MiddleEar_TympanicCavity_L': 20, 'MiddleEar_TympanicCavity_R': 21, 'TympanicCavity_L': 22, 'TympanicCavity_R': 23, 'MiddleEar_VestibulSemi_L': 24, 'MiddleEar_VestibulSemi_R': 25, 'VestibulSemi_L': 26, 'VestibulSemi_R': 27, 'Cochlea_L': 28, 'Cochlea_R': 29, 'MiddleEar_ETbone_L': 30, 'MiddleEar_ETbone_R': 31, 'ETbone_L': 32, 'ETbone_R': 33, 'Pituitary': 34, 'OralCavity': 35, 'Mandible_L': 36, 'Mandible_R': 37, 'Submandibular_L': 38, 'Submandibular_R': 39, 'Parotid_L': 40, 'Parotid_R': 41, 'Mastoid_L': 42, 'Mastoid_R': 43, 'TMjoint_L': 44, 'TMjoint_R': 45, 'SpinalCord': 46, 'Esophagus': 47, 'Larynx': 48, 'Larynx_Glottic': 49, 'Larynx_Supraglot': 50, 'Larynx_PharynxConst': 51, 'PharynxConst': 52, 'Thyroid': 53, 'Trachea': 54}
OAR_TO_LABEL_IDS = {'Brain': [1, 2, 3, 4, 5, 6, 7, 8, 9], 'BrainStem': [2], 'Chiasm': [3], 'TemporalLobe_L': [4, 6], 'TemporalLobe_R': [5, 7], 'Hippocampus_L': [8, 6], 'Hippocampus_R': [9, 7], 'Eye_L': [10, 12], 'Eye_R': [11, 13], 'Lens_L': [12], 'Lens_R': [13], 'OpticNerve_L': [14], 'OpticNerve_R': [15], 'MiddleEar_L': [18, 16, 20, 24, 28, 30], 'MiddleEar_R': [19, 17, 21, 25, 29, 31], 'IAC_L': [18], 'IAC_R': [19], 'TympanicCavity_L': [22, 20], 'TympanicCavity_R': [23, 21], 'VestibulSemi_L': [26, 24], 'VestibulSemi_R': [27, 25], 'Cochlea_L': [28], 'Cochlea_R': [29], 'ETbone_L': [32, 30], 'ETbone_R': [33, 31], 'Pituitary': [34], 'OralCavity': [35], 'Mandible_L': [36], 'Mandible_R': [37], 'Submandibular_L': [38], 'Submandibular_R': [39], 'Parotid_L': [40], 'Parotid_R': [41], 'Mastoid_L': [42], 'Mastoid_R': [43], 'TMjoint_L': [44], 'TMjoint_R': [45], 'SpinalCord': [46], 'Esophagus': [47], 'Larynx': [48, 49, 50, 51], 'Larynx_Glottic': [49], 'Larynx_Supraglot': [50], 'PharynxConst': [51, 52], 'Thyroid': [53], 'Trachea': [54]}
def get_segrap_data( path: Union[os.PathLike, str], modality: Literal['ct', 'ct_contrast'] = 'ct', download: bool = False, task: Literal['oars', 'gtv'] = 'oars') -> str:
165def get_segrap_data(
166    path: Union[os.PathLike, str],
167    modality: Literal["ct", "ct_contrast"] = "ct",
168    download: bool = False,
169    task: Literal["oars", "gtv"] = "oars",
170) -> str:
171    """Download the SegRap dataset.
172
173    Args:
174        path: Filepath to a folder where the data is downloaded for further processing.
175        modality: The CT scan to download. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
176        download: Whether to download the data if it is not present.
177        task: The annotations to download. Either 'oars' (Task001) or 'gtv' (Task002).
178
179    Returns:
180        Filepath where the data is stored.
181    """
182    if modality not in OFFICIAL_FILE_NAMES:
183        raise ValueError(f"'{modality}' is not a valid modality. Choose one of {list(OFFICIAL_FILE_NAMES)}.")
184    if task not in TASKS:
185        raise ValueError(f"'{task}' is not a valid task. Choose one of {list(TASKS)}.")
186
187    if len(_get_official_case_dirs(path)) > 0:  # The official data was downloaded manually.
188        return path
189
190    _download_component(path, TASKS[task][0], download)
191    _download_component(path, modality, download)
192
193    return path

Download the SegRap dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • modality: The CT scan to download. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • download: Whether to download the data if it is not present.
  • task: The annotations to download. Either 'oars' (Task001) or 'gtv' (Task002).
Returns:

Filepath where the data is stored.

def get_segrap_paths( path: Union[os.PathLike, str], modality: Literal['ct', 'ct_contrast'] = 'ct', download: bool = False, task: Literal['oars', 'gtv'] = 'oars') -> Tuple[List[str], List[str]]:
196def get_segrap_paths(
197    path: Union[os.PathLike, str],
198    modality: Literal["ct", "ct_contrast"] = "ct",
199    download: bool = False,
200    task: Literal["oars", "gtv"] = "oars",
201) -> Tuple[List[str], List[str]]:
202    """Get paths to the SegRap data.
203
204    Args:
205        path: Filepath to a folder where the data is downloaded for further processing.
206        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
207        download: Whether to download the data if it is not present.
208        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
209
210    Returns:
211        List of filepaths for the image data.
212        List of filepaths for the label data.
213    """
214    data_dir = get_segrap_data(path, modality, download, task)
215    label_component, official_label_dir = TASKS[task]
216
217    case_dirs = _get_official_case_dirs(data_dir)
218    if len(case_dirs) > 0:  # The official layout, with one folder per case and the labels in a separate folder.
219        label_dir = os.path.join(os.path.split(os.path.split(case_dirs[0])[0])[0], official_label_dir)
220        raw_paths = [os.path.join(p, OFFICIAL_FILE_NAMES[modality]) for p in case_dirs]
221        label_paths = [os.path.join(label_dir, f"{os.path.basename(p)}.nii.gz") for p in case_dirs]
222    else:  # The redistributed layout, with the files named after the case ids.
223        raw_paths = natsorted(glob(os.path.join(data_dir, FOLDER_NAMES[modality], "*.nii.gz")))
224        label_paths = [
225            os.path.join(data_dir, FOLDER_NAMES[label_component], os.path.basename(p)) for p in raw_paths
226        ]
227
228    assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths + label_paths)
229
230    return raw_paths, label_paths

Get paths to the SegRap data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • download: Whether to download the data if it is not present.
  • task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_segrap_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], modality: Literal['ct', 'ct_contrast'] = 'ct', resize_inputs: bool = False, download: bool = False, task: Literal['oars', 'gtv'] = 'oars', **kwargs) -> torch.utils.data.dataset.Dataset:
233def get_segrap_dataset(
234    path: Union[os.PathLike, str],
235    patch_shape: Tuple[int, ...],
236    modality: Literal["ct", "ct_contrast"] = "ct",
237    resize_inputs: bool = False,
238    download: bool = False,
239    task: Literal["oars", "gtv"] = "oars",
240    **kwargs
241) -> Dataset:
242    """Get the SegRap dataset for organ-at-risk or GTV segmentation.
243
244    Args:
245        path: Filepath to a folder where the data is downloaded for further processing.
246        patch_shape: The patch shape to use for training.
247        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
248        resize_inputs: Whether to resize inputs to the desired patch shape.
249        download: Whether to download the data if it is not present.
250        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
251        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
252
253    Returns:
254        The segmentation dataset.
255    """
256    raw_paths, label_paths = get_segrap_paths(path, modality, download, task)
257
258    if resize_inputs:
259        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
260        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
261            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
262        )
263
264    return torch_em.default_segmentation_dataset(
265        raw_paths=raw_paths,
266        raw_key="data",
267        label_paths=label_paths,
268        label_key="data",
269        patch_shape=patch_shape,
270        is_seg_dataset=True,
271        **kwargs
272    )

Get the SegRap dataset for organ-at-risk or GTV segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_segrap_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], modality: Literal['ct', 'ct_contrast'] = 'ct', resize_inputs: bool = False, download: bool = False, task: Literal['oars', 'gtv'] = 'oars', **kwargs) -> torch.utils.data.dataloader.DataLoader:
275def get_segrap_loader(
276    path: Union[os.PathLike, str],
277    batch_size: int,
278    patch_shape: Tuple[int, ...],
279    modality: Literal["ct", "ct_contrast"] = "ct",
280    resize_inputs: bool = False,
281    download: bool = False,
282    task: Literal["oars", "gtv"] = "oars",
283    **kwargs
284) -> DataLoader:
285    """Get the SegRap dataloader for organ-at-risk or GTV segmentation.
286
287    Args:
288        path: Filepath to a folder where the data is downloaded for further processing.
289        batch_size: The batch size for training.
290        patch_shape: The patch shape to use for training.
291        modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
292        resize_inputs: Whether to resize inputs to the desired patch shape.
293        download: Whether to download the data if it is not present.
294        task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
295        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
296
297    Returns:
298        The DataLoader.
299    """
300    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
301    dataset = get_segrap_dataset(path, patch_shape, modality, resize_inputs, download, task, **ds_kwargs)
302    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the SegRap dataloader for organ-at-risk or GTV segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • modality: The CT scan to use as input. Either 'ct' (non-contrast) or 'ct_contrast' (contrast-enhanced).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • task: The annotations to use. Either 'oars' (Task001) or 'gtv' (Task002).
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.