torch_em.data.datasets.medical.nsclc_radiomics_interobserver1
The NSCLC-Radiomics-Interobserver1 dataset contains annotations for primary gross tumor volume (GTV) segmentation in preoperative CT of non-small cell lung cancer patients, from an interobserver-variability study.
It consists of 22 CT volumes, 21 of which have a primary GTV delineated independently by 5 radiation oncologists, distributed as DICOM-SEG objects and converted and stored in hdf5 files by this module. Each radiation oncologist provided two delineations of the same tumor: a purely manual ('vis') one and one assisted by an autosegmentation tool and then manually edited ('auto'). Radiation oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. The semantic label id is: 1: tumor. This module queries the TCIA REST API for the segmented series and downloads only those and their referenced CT series, rather than the full collection manifest (which also has the CT series of the one patient without a usable delineation).
NOTE: This requires the pydicom python package.
The dataset is located at https://www.cancerimagingarchive.net/collection/nsclc-radiomics-interobserver1/. It is released under the CC BY-NC 3.0 license.
This dataset is from the publication https://doi.org/10.1038/ncomms5006. The data was released at https://doi.org/10.7937/tcia.2019.cwvlpd26. Please cite it if you use this dataset in your research.
1"""The NSCLC-Radiomics-Interobserver1 dataset contains annotations for primary gross tumor volume (GTV) 2segmentation in preoperative CT of non-small cell lung cancer patients, from an interobserver-variability 3study. 4 5It consists of 22 CT volumes, 21 of which have a primary GTV delineated independently by 5 radiation 6oncologists, distributed as DICOM-SEG objects and converted and stored in hdf5 files by this module. Each 7radiation oncologist provided two delineations of the same tumor: a purely manual ('vis') one and one 8assisted by an autosegmentation tool and then manually edited ('auto'). Radiation oncologists '1' and '3' 9were trainees at the time of the study, '2', '4' and '5' were extensively experienced. The semantic label 10id is: 1: tumor. This module queries the TCIA REST API for the segmented series and downloads only those 11and their referenced CT series, rather than the full collection manifest (which also has the CT series 12of the one patient without a usable delineation). 13 14NOTE: This requires the pydicom python package. 15 16The dataset is located at https://www.cancerimagingarchive.net/collection/nsclc-radiomics-interobserver1/. 17It is released under the CC BY-NC 3.0 license. 18 19This dataset is from the publication https://doi.org/10.1038/ncomms5006. 20The data was released at https://doi.org/10.7937/tcia.2019.cwvlpd26. 21Please cite it if you use this dataset in your research. 22""" 23 24import os 25import csv 26from glob import glob 27from tqdm import tqdm 28from warnings import warn 29from natsort import natsorted 30from collections import defaultdict 31from typing import Union, Tuple, List, Literal 32 33import requests 34import numpy as np 35 36from torch.utils.data import Dataset, DataLoader 37 38import torch_em 39 40from .. import util 41 42 43COLLECTION = "NSCLC-Radiomics-Interobserver1" 44 45LABEL_IDS = {"tumor": 1} 46 47ANNOTATORS = (1, 2, 3, 4, 5) 48 49# The DICOM-SEG segment descriptions follow the pattern 'GTV-1<annotation_type><annotator>', e.g. 50# 'GTV-1vis-3' for the manual delineation of annotator 3, see the module docstring. 51_ANNOTATION_TYPE_TO_TAG = {"manual": "vis", "auto": "auto"} 52 53 54def _get_segmented_series_uids(): 55 """Query the TCIA REST API for the SEG series of the collection, without downloading any image data.""" 56 response = requests.get(util.NBIA_API_URL + "getSeries", params={"Collection": COLLECTION, "Modality": "SEG"}) 57 response.raise_for_status() 58 return [row["SeriesInstanceUID"] for row in response.json()] 59 60 61def _get_referenced_sop_uids(seg_path): 62 """Get the SOP instance UIDs of the CT slices referenced by a DICOM-SEG object.""" 63 import pydicom 64 65 seg = pydicom.dcmread(seg_path, stop_before_pixels=True) 66 assert len(seg.ReferencedSeriesSequence) == 1, f"Expected a single referenced CT series in {seg_path}." 67 referenced_instances = seg.ReferencedSeriesSequence[0].ReferencedInstanceSequence 68 return {str(instance.ReferencedSOPInstanceUID) for instance in referenced_instances} 69 70 71def _get_referenced_ct_series_uid(seg_path): 72 """Get the series instance UID of the CT series referenced by a DICOM-SEG object.""" 73 import pydicom 74 75 seg = pydicom.dcmread(seg_path, stop_before_pixels=True) 76 assert len(seg.ReferencedSeriesSequence) == 1, f"Expected a single referenced CT series in {seg_path}." 77 return str(seg.ReferencedSeriesSequence[0].SeriesInstanceUID) 78 79 80def _load_dicom_volume(series_dir, referenced_sop_uids): 81 """Stack a DICOM series into a volume with axes (z, y, x), sorted by ascending patient z position. 82 83 Returns the volume in Hounsfield units and the geometry needed to align the DICOM-SEG frames with the volume. 84 """ 85 import pydicom 86 87 slices = [pydicom.dcmread(dcm_path) for dcm_path in natsorted(glob(os.path.join(series_dir, "*.dcm")))] 88 slices = [dcm for dcm in slices if dcm.SOPInstanceUID in referenced_sop_uids] 89 slices.sort(key=lambda dcm: dcm.ImagePositionPatient[2]) 90 91 volume = np.stack([dcm.pixel_array for dcm in slices]).astype("float32") 92 volume = volume * float(slices[0].RescaleSlope) + float(slices[0].RescaleIntercept) 93 volume = np.round(volume).astype("int16") 94 95 geometry = { 96 "sop_uids": {str(dcm.SOPInstanceUID): i for i, dcm in enumerate(slices)}, 97 "z_positions": np.array([float(dcm.ImagePositionPatient[2]) for dcm in slices]), 98 "orientation": np.round([float(v) for v in slices[0].ImageOrientationPatient]).astype("int"), 99 } 100 return volume, geometry 101 102 103def _parse_segment_description(description): 104 """Parse a DICOM-SEG segment description, e.g. 'GTV-1vis-3', into an (annotation_type, annotator) key.""" 105 for annotation_type, tag in _ANNOTATION_TYPE_TO_TAG.items(): 106 prefix = f"GTV-1{tag}-" 107 if str(description).startswith(prefix): 108 annotator = str(description)[len(prefix):] 109 if annotator in {str(a) for a in ANNOTATORS}: 110 return annotation_type, annotator 111 return None 112 113 114def _load_dicom_seg(seg_path, shape, geometry): 115 """Convert a DICOM-SEG object into per-annotator, per-annotation-type binary tumor masks. 116 117 Each frame is mapped to its CT slice via the source image it was derived from (or its z position), and 118 to a segment via its 'ReferencedSegmentNumber'. Only the segments matching the 'GTV-1<vis|auto>-<1-5>' 119 naming convention are kept (see the module docstring); the additional PET-SUV-threshold-based ROIs that 120 some patients also carry are ignored, since they are not tied to a specific annotator. 121 122 Returns a dict mapping '<annotation_type>/<annotator>' (e.g. 'manual/3') to a boolean mask. 123 """ 124 import pydicom 125 126 seg = pydicom.dcmread(seg_path) 127 128 segment_keys = {} 129 for segment in seg.SegmentSequence: 130 parsed = _parse_segment_description(segment.SegmentDescription) 131 if parsed is not None: 132 annotation_type, annotator = parsed 133 segment_keys[int(segment.SegmentNumber)] = f"{annotation_type}/{annotator}" 134 135 seg_orientation = seg.SharedFunctionalGroupsSequence[0].PlaneOrientationSequence[0].ImageOrientationPatient 136 seg_orientation = np.round([float(v) for v in seg_orientation]).astype("int") 137 orientation = geometry["orientation"] 138 assert np.all(np.abs(seg_orientation) == np.abs(orientation)), f"Unexpected orientation in {seg_path}." 139 flip_axes = [] 140 if np.any(seg_orientation[3:] != orientation[3:]): # The direction of the rows differs. 141 flip_axes.append(0) 142 if np.any(seg_orientation[:3] != orientation[:3]): # The direction of the columns differs. 143 flip_axes.append(1) 144 145 z_positions = geometry["z_positions"] 146 tolerance = np.diff(z_positions).min() / 2 if len(z_positions) > 1 else 1.0 147 148 frames = seg.pixel_array 149 if frames.ndim == 2: # A segmentation with a single frame. 150 frames = frames[None] 151 152 masks = {key: np.zeros(shape, dtype="bool") for key in segment_keys.values()} 153 for frame, frame_group in zip(frames, seg.PerFrameFunctionalGroupsSequence): 154 segment_number = int(frame_group.SegmentIdentificationSequence[0].ReferencedSegmentNumber) 155 key = segment_keys.get(segment_number) 156 if key is None: # Not one of the 'GTV-1<vis|auto>-<1-5>' segments, e.g. a PET-SUV-threshold-based ROI. 157 continue 158 frame_mask = frame.astype("bool") 159 160 z = None 161 derivation = frame_group.get("DerivationImageSequence", []) 162 if derivation and derivation[0].get("SourceImageSequence"): 163 z = geometry["sop_uids"].get(str(derivation[0].SourceImageSequence[0].ReferencedSOPInstanceUID)) 164 if z is None: 165 frame_z = float(frame_group.PlanePositionSequence[0].ImagePositionPatient[2]) 166 z = int(np.argmin(np.abs(z_positions - frame_z))) 167 if abs(z_positions[z] - frame_z) > tolerance: 168 if frame_mask.any(): 169 warn(f"Skipping a frame at z={frame_z} in {seg_path}, which does not match a CT slice.") 170 continue 171 172 if flip_axes: 173 frame_mask = np.flip(frame_mask, axis=flip_axes) 174 masks[key][z] |= frame_mask 175 176 return masks 177 178 179def _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir): 180 import h5py 181 182 series_per_subject = defaultdict(dict) 183 for csv_path in csv_paths: 184 with open(csv_path, "r") as f: 185 for row in csv.DictReader(f): 186 series_per_subject[row["Subject ID"]][row["Modality"]] = os.path.join(dicom_dir, row["Series UID"]) 187 188 os.makedirs(preprocessed_dir, exist_ok=True) 189 subjects_with_seg = {sid: series for sid, series in series_per_subject.items() if "SEG" in series} 190 for subject_id, series_dirs in tqdm( 191 sorted(subjects_with_seg.items()), desc="Preprocess NSCLC-Radiomics-Interobserver1" 192 ): 193 out_path = os.path.join(preprocessed_dir, f"{subject_id}.h5") 194 if os.path.exists(out_path): 195 continue 196 if "CT" not in series_dirs: 197 warn(f"Skipping {subject_id}, which has a SEG object but no CT series.") 198 continue 199 200 seg_path = glob(os.path.join(series_dirs["SEG"], "*.dcm"))[0] 201 volume, geometry = _load_dicom_volume(series_dirs["CT"], _get_referenced_sop_uids(seg_path)) 202 masks = _load_dicom_seg(seg_path, volume.shape, geometry) 203 204 with h5py.File(out_path, "w") as f: 205 f.create_dataset("raw", data=volume, compression="gzip") 206 for key, mask in masks.items(): 207 labels = (mask * LABEL_IDS["tumor"]).astype("uint8") 208 f.create_dataset(f"labels/{key}", data=labels, compression="gzip") 209 210 211def get_nsclc_radiomics_interobserver1_data(path: Union[os.PathLike, str], download: bool = False) -> str: 212 """Download the NSCLC-Radiomics-Interobserver1 dataset. 213 214 Args: 215 path: Filepath to a folder where the data is downloaded for further processing. 216 download: Whether to download the data if it is not present. 217 218 Returns: 219 Filepath where the preprocessed data is stored. 220 """ 221 # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes. 222 preprocessed_dir = os.path.join(path, "preprocessed") 223 os.makedirs(path, exist_ok=True) 224 225 # Only the segmented series and their referenced CT series are downloaded (see the module docstring), 226 # not the full collection manifest. Each download step is skipped once its metadata csv is written. 227 dicom_dir = os.path.join(path, "dicom") 228 seg_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_seg_series") 229 if not os.path.exists(f"{seg_csv_path}.csv"): 230 if not download: 231 raise RuntimeError(f"Cannot find the data at {dicom_dir}, but download was set to False.") 232 seg_uids = _get_segmented_series_uids() 233 util.download_tcia_series(seg_uids, dicom_dir, seg_csv_path) 234 235 ct_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_ct_series") 236 if not os.path.exists(f"{ct_csv_path}.csv"): 237 with open(f"{seg_csv_path}.csv", "r") as f: 238 seg_uids = [row["Series UID"] for row in csv.DictReader(f)] 239 ct_uids = sorted({ 240 _get_referenced_ct_series_uid(glob(os.path.join(dicom_dir, uid, "*.dcm"))[0]) for uid in seg_uids 241 }) 242 util.download_tcia_series(ct_uids, dicom_dir, ct_csv_path) 243 244 csv_paths = [f"{seg_csv_path}.csv", f"{ct_csv_path}.csv"] 245 _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir) 246 return preprocessed_dir 247 248 249def get_nsclc_radiomics_interobserver1_paths( 250 path: Union[os.PathLike, str], 251 annotator: int = 1, 252 annotation_type: Literal["manual", "auto"] = "manual", 253 download: bool = False, 254) -> Tuple[List[str], str]: 255 """Get paths to the NSCLC-Radiomics-Interobserver1 data. 256 257 Args: 258 path: Filepath to a folder where the data is downloaded for further processing. 259 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 260 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 261 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 262 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 263 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 264 download: Whether to download the data if it is not present. 265 266 Returns: 267 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data. 268 The key of the label data in each hdf5 file. 269 """ 270 assert annotator in ANNOTATORS, f"Invalid annotator: {annotator}. Choose from {ANNOTATORS}." 271 assert annotation_type in _ANNOTATION_TYPE_TO_TAG, f"Invalid annotation_type: {annotation_type}." 272 273 import h5py 274 275 data_dir = get_nsclc_radiomics_interobserver1_data(path, download) 276 label_key = f"labels/{annotation_type}/{annotator}" 277 278 volume_paths = [] 279 for volume_path in natsorted(glob(os.path.join(data_dir, "*.h5"))): 280 with h5py.File(volume_path, "r") as f: 281 has_label = label_key in f 282 if has_label: 283 volume_paths.append(volume_path) 284 else: 285 warn(f"Skipping {volume_path}, which has no '{label_key}' delineation.") 286 287 return volume_paths, label_key 288 289 290def get_nsclc_radiomics_interobserver1_dataset( 291 path: Union[os.PathLike, str], 292 patch_shape: Tuple[int, ...], 293 annotator: int = 1, 294 annotation_type: Literal["manual", "auto"] = "manual", 295 resize_inputs: bool = False, 296 download: bool = False, 297 **kwargs 298) -> Dataset: 299 """Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation. 300 301 Args: 302 path: Filepath to a folder where the data is downloaded for further processing. 303 patch_shape: The patch shape to use for training. 304 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 305 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 306 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 307 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 308 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 309 resize_inputs: Whether to resize inputs to the desired patch shape. 310 download: Whether to download the data if it is not present. 311 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 312 313 Returns: 314 The segmentation dataset. 315 """ 316 volume_paths, label_key = get_nsclc_radiomics_interobserver1_paths(path, annotator, annotation_type, download) 317 318 if resize_inputs: 319 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 320 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 321 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 322 ) 323 324 return torch_em.default_segmentation_dataset( 325 raw_paths=volume_paths, 326 raw_key="raw", 327 label_paths=volume_paths, 328 label_key=label_key, 329 patch_shape=patch_shape, 330 is_seg_dataset=True, 331 **kwargs 332 ) 333 334 335def get_nsclc_radiomics_interobserver1_loader( 336 path: Union[os.PathLike, str], 337 batch_size: int, 338 patch_shape: Tuple[int, ...], 339 annotator: int = 1, 340 annotation_type: Literal["manual", "auto"] = "manual", 341 resize_inputs: bool = False, 342 download: bool = False, 343 **kwargs 344) -> DataLoader: 345 """Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation. 346 347 Args: 348 path: Filepath to a folder where the data is downloaded for further processing. 349 batch_size: The batch size for training. 350 patch_shape: The patch shape to use for training. 351 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 352 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 353 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 354 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 355 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 356 resize_inputs: Whether to resize inputs to the desired patch shape. 357 download: Whether to download the data if it is not present. 358 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 359 360 Returns: 361 The DataLoader. 362 """ 363 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 364 dataset = get_nsclc_radiomics_interobserver1_dataset( 365 path, patch_shape, annotator, annotation_type, resize_inputs, download, **ds_kwargs 366 ) 367 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
212def get_nsclc_radiomics_interobserver1_data(path: Union[os.PathLike, str], download: bool = False) -> str: 213 """Download the NSCLC-Radiomics-Interobserver1 dataset. 214 215 Args: 216 path: Filepath to a folder where the data is downloaded for further processing. 217 download: Whether to download the data if it is not present. 218 219 Returns: 220 Filepath where the preprocessed data is stored. 221 """ 222 # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes. 223 preprocessed_dir = os.path.join(path, "preprocessed") 224 os.makedirs(path, exist_ok=True) 225 226 # Only the segmented series and their referenced CT series are downloaded (see the module docstring), 227 # not the full collection manifest. Each download step is skipped once its metadata csv is written. 228 dicom_dir = os.path.join(path, "dicom") 229 seg_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_seg_series") 230 if not os.path.exists(f"{seg_csv_path}.csv"): 231 if not download: 232 raise RuntimeError(f"Cannot find the data at {dicom_dir}, but download was set to False.") 233 seg_uids = _get_segmented_series_uids() 234 util.download_tcia_series(seg_uids, dicom_dir, seg_csv_path) 235 236 ct_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_ct_series") 237 if not os.path.exists(f"{ct_csv_path}.csv"): 238 with open(f"{seg_csv_path}.csv", "r") as f: 239 seg_uids = [row["Series UID"] for row in csv.DictReader(f)] 240 ct_uids = sorted({ 241 _get_referenced_ct_series_uid(glob(os.path.join(dicom_dir, uid, "*.dcm"))[0]) for uid in seg_uids 242 }) 243 util.download_tcia_series(ct_uids, dicom_dir, ct_csv_path) 244 245 csv_paths = [f"{seg_csv_path}.csv", f"{ct_csv_path}.csv"] 246 _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir) 247 return preprocessed_dir
Download the NSCLC-Radiomics-Interobserver1 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the preprocessed data is stored.
250def get_nsclc_radiomics_interobserver1_paths( 251 path: Union[os.PathLike, str], 252 annotator: int = 1, 253 annotation_type: Literal["manual", "auto"] = "manual", 254 download: bool = False, 255) -> Tuple[List[str], str]: 256 """Get paths to the NSCLC-Radiomics-Interobserver1 data. 257 258 Args: 259 path: Filepath to a folder where the data is downloaded for further processing. 260 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 261 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 262 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 263 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 264 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 265 download: Whether to download the data if it is not present. 266 267 Returns: 268 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data. 269 The key of the label data in each hdf5 file. 270 """ 271 assert annotator in ANNOTATORS, f"Invalid annotator: {annotator}. Choose from {ANNOTATORS}." 272 assert annotation_type in _ANNOTATION_TYPE_TO_TAG, f"Invalid annotation_type: {annotation_type}." 273 274 import h5py 275 276 data_dir = get_nsclc_radiomics_interobserver1_data(path, download) 277 label_key = f"labels/{annotation_type}/{annotator}" 278 279 volume_paths = [] 280 for volume_path in natsorted(glob(os.path.join(data_dir, "*.h5"))): 281 with h5py.File(volume_path, "r") as f: 282 has_label = label_key in f 283 if has_label: 284 volume_paths.append(volume_path) 285 else: 286 warn(f"Skipping {volume_path}, which has no '{label_key}' delineation.") 287 288 return volume_paths, label_key
Get paths to the NSCLC-Radiomics-Interobserver1 data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
- annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data. The key of the label data in each hdf5 file.
291def get_nsclc_radiomics_interobserver1_dataset( 292 path: Union[os.PathLike, str], 293 patch_shape: Tuple[int, ...], 294 annotator: int = 1, 295 annotation_type: Literal["manual", "auto"] = "manual", 296 resize_inputs: bool = False, 297 download: bool = False, 298 **kwargs 299) -> Dataset: 300 """Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation. 301 302 Args: 303 path: Filepath to a folder where the data is downloaded for further processing. 304 patch_shape: The patch shape to use for training. 305 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 306 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 307 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 308 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 309 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 310 resize_inputs: Whether to resize inputs to the desired patch shape. 311 download: Whether to download the data if it is not present. 312 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 313 314 Returns: 315 The segmentation dataset. 316 """ 317 volume_paths, label_key = get_nsclc_radiomics_interobserver1_paths(path, annotator, annotation_type, download) 318 319 if resize_inputs: 320 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 321 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 322 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 323 ) 324 325 return torch_em.default_segmentation_dataset( 326 raw_paths=volume_paths, 327 raw_key="raw", 328 label_paths=volume_paths, 329 label_key=label_key, 330 patch_shape=patch_shape, 331 is_seg_dataset=True, 332 **kwargs 333 )
Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
- annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
336def get_nsclc_radiomics_interobserver1_loader( 337 path: Union[os.PathLike, str], 338 batch_size: int, 339 patch_shape: Tuple[int, ...], 340 annotator: int = 1, 341 annotation_type: Literal["manual", "auto"] = "manual", 342 resize_inputs: bool = False, 343 download: bool = False, 344 **kwargs 345) -> DataLoader: 346 """Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation. 347 348 Args: 349 path: Filepath to a folder where the data is downloaded for further processing. 350 batch_size: The batch size for training. 351 patch_shape: The patch shape to use for training. 352 annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' 353 and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. 354 annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 355 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of 356 the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice. 357 resize_inputs: Whether to resize inputs to the desired patch shape. 358 download: Whether to download the data if it is not present. 359 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 360 361 Returns: 362 The DataLoader. 363 """ 364 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 365 dataset = get_nsclc_radiomics_interobserver1_dataset( 366 path, patch_shape, annotator, annotation_type, resize_inputs, download, **ds_kwargs 367 ) 368 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
- annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.