torch_em.data.datasets.medical.nsclc_radiomics_interobserver1

The NSCLC-Radiomics-Interobserver1 dataset contains annotations for primary gross tumor volume (GTV) segmentation in preoperative CT of non-small cell lung cancer patients, from an interobserver-variability study.

It consists of 22 CT volumes, 21 of which have a primary GTV delineated independently by 5 radiation oncologists, distributed as DICOM-SEG objects and converted and stored in hdf5 files by this module. Each radiation oncologist provided two delineations of the same tumor: a purely manual ('vis') one and one assisted by an autosegmentation tool and then manually edited ('auto'). Radiation oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced. The semantic label id is: 1: tumor. This module queries the TCIA REST API for the segmented series and downloads only those and their referenced CT series, rather than the full collection manifest (which also has the CT series of the one patient without a usable delineation).

NOTE: This requires the pydicom python package.

The dataset is located at https://www.cancerimagingarchive.net/collection/nsclc-radiomics-interobserver1/. It is released under the CC BY-NC 3.0 license.

This dataset is from the publication https://doi.org/10.1038/ncomms5006. The data was released at https://doi.org/10.7937/tcia.2019.cwvlpd26. Please cite it if you use this dataset in your research.

  1"""The NSCLC-Radiomics-Interobserver1 dataset contains annotations for primary gross tumor volume (GTV)
  2segmentation in preoperative CT of non-small cell lung cancer patients, from an interobserver-variability
  3study.
  4
  5It consists of 22 CT volumes, 21 of which have a primary GTV delineated independently by 5 radiation
  6oncologists, distributed as DICOM-SEG objects and converted and stored in hdf5 files by this module. Each
  7radiation oncologist provided two delineations of the same tumor: a purely manual ('vis') one and one
  8assisted by an autosegmentation tool and then manually edited ('auto'). Radiation oncologists '1' and '3'
  9were trainees at the time of the study, '2', '4' and '5' were extensively experienced. The semantic label
 10id is: 1: tumor. This module queries the TCIA REST API for the segmented series and downloads only those
 11and their referenced CT series, rather than the full collection manifest (which also has the CT series
 12of the one patient without a usable delineation).
 13
 14NOTE: This requires the pydicom python package.
 15
 16The dataset is located at https://www.cancerimagingarchive.net/collection/nsclc-radiomics-interobserver1/.
 17It is released under the CC BY-NC 3.0 license.
 18
 19This dataset is from the publication https://doi.org/10.1038/ncomms5006.
 20The data was released at https://doi.org/10.7937/tcia.2019.cwvlpd26.
 21Please cite it if you use this dataset in your research.
 22"""
 23
 24import os
 25import csv
 26from glob import glob
 27from tqdm import tqdm
 28from warnings import warn
 29from natsort import natsorted
 30from collections import defaultdict
 31from typing import Union, Tuple, List, Literal
 32
 33import requests
 34import numpy as np
 35
 36from torch.utils.data import Dataset, DataLoader
 37
 38import torch_em
 39
 40from .. import util
 41
 42
 43COLLECTION = "NSCLC-Radiomics-Interobserver1"
 44
 45LABEL_IDS = {"tumor": 1}
 46
 47ANNOTATORS = (1, 2, 3, 4, 5)
 48
 49# The DICOM-SEG segment descriptions follow the pattern 'GTV-1<annotation_type><annotator>', e.g.
 50# 'GTV-1vis-3' for the manual delineation of annotator 3, see the module docstring.
 51_ANNOTATION_TYPE_TO_TAG = {"manual": "vis", "auto": "auto"}
 52
 53
 54def _get_segmented_series_uids():
 55    """Query the TCIA REST API for the SEG series of the collection, without downloading any image data."""
 56    response = requests.get(util.NBIA_API_URL + "getSeries", params={"Collection": COLLECTION, "Modality": "SEG"})
 57    response.raise_for_status()
 58    return [row["SeriesInstanceUID"] for row in response.json()]
 59
 60
 61def _get_referenced_sop_uids(seg_path):
 62    """Get the SOP instance UIDs of the CT slices referenced by a DICOM-SEG object."""
 63    import pydicom
 64
 65    seg = pydicom.dcmread(seg_path, stop_before_pixels=True)
 66    assert len(seg.ReferencedSeriesSequence) == 1, f"Expected a single referenced CT series in {seg_path}."
 67    referenced_instances = seg.ReferencedSeriesSequence[0].ReferencedInstanceSequence
 68    return {str(instance.ReferencedSOPInstanceUID) for instance in referenced_instances}
 69
 70
 71def _get_referenced_ct_series_uid(seg_path):
 72    """Get the series instance UID of the CT series referenced by a DICOM-SEG object."""
 73    import pydicom
 74
 75    seg = pydicom.dcmread(seg_path, stop_before_pixels=True)
 76    assert len(seg.ReferencedSeriesSequence) == 1, f"Expected a single referenced CT series in {seg_path}."
 77    return str(seg.ReferencedSeriesSequence[0].SeriesInstanceUID)
 78
 79
 80def _load_dicom_volume(series_dir, referenced_sop_uids):
 81    """Stack a DICOM series into a volume with axes (z, y, x), sorted by ascending patient z position.
 82
 83    Returns the volume in Hounsfield units and the geometry needed to align the DICOM-SEG frames with the volume.
 84    """
 85    import pydicom
 86
 87    slices = [pydicom.dcmread(dcm_path) for dcm_path in natsorted(glob(os.path.join(series_dir, "*.dcm")))]
 88    slices = [dcm for dcm in slices if dcm.SOPInstanceUID in referenced_sop_uids]
 89    slices.sort(key=lambda dcm: dcm.ImagePositionPatient[2])
 90
 91    volume = np.stack([dcm.pixel_array for dcm in slices]).astype("float32")
 92    volume = volume * float(slices[0].RescaleSlope) + float(slices[0].RescaleIntercept)
 93    volume = np.round(volume).astype("int16")
 94
 95    geometry = {
 96        "sop_uids": {str(dcm.SOPInstanceUID): i for i, dcm in enumerate(slices)},
 97        "z_positions": np.array([float(dcm.ImagePositionPatient[2]) for dcm in slices]),
 98        "orientation": np.round([float(v) for v in slices[0].ImageOrientationPatient]).astype("int"),
 99    }
100    return volume, geometry
101
102
103def _parse_segment_description(description):
104    """Parse a DICOM-SEG segment description, e.g. 'GTV-1vis-3', into an (annotation_type, annotator) key."""
105    for annotation_type, tag in _ANNOTATION_TYPE_TO_TAG.items():
106        prefix = f"GTV-1{tag}-"
107        if str(description).startswith(prefix):
108            annotator = str(description)[len(prefix):]
109            if annotator in {str(a) for a in ANNOTATORS}:
110                return annotation_type, annotator
111    return None
112
113
114def _load_dicom_seg(seg_path, shape, geometry):
115    """Convert a DICOM-SEG object into per-annotator, per-annotation-type binary tumor masks.
116
117    Each frame is mapped to its CT slice via the source image it was derived from (or its z position), and
118    to a segment via its 'ReferencedSegmentNumber'. Only the segments matching the 'GTV-1<vis|auto>-<1-5>'
119    naming convention are kept (see the module docstring); the additional PET-SUV-threshold-based ROIs that
120    some patients also carry are ignored, since they are not tied to a specific annotator.
121
122    Returns a dict mapping '<annotation_type>/<annotator>' (e.g. 'manual/3') to a boolean mask.
123    """
124    import pydicom
125
126    seg = pydicom.dcmread(seg_path)
127
128    segment_keys = {}
129    for segment in seg.SegmentSequence:
130        parsed = _parse_segment_description(segment.SegmentDescription)
131        if parsed is not None:
132            annotation_type, annotator = parsed
133            segment_keys[int(segment.SegmentNumber)] = f"{annotation_type}/{annotator}"
134
135    seg_orientation = seg.SharedFunctionalGroupsSequence[0].PlaneOrientationSequence[0].ImageOrientationPatient
136    seg_orientation = np.round([float(v) for v in seg_orientation]).astype("int")
137    orientation = geometry["orientation"]
138    assert np.all(np.abs(seg_orientation) == np.abs(orientation)), f"Unexpected orientation in {seg_path}."
139    flip_axes = []
140    if np.any(seg_orientation[3:] != orientation[3:]):  # The direction of the rows differs.
141        flip_axes.append(0)
142    if np.any(seg_orientation[:3] != orientation[:3]):  # The direction of the columns differs.
143        flip_axes.append(1)
144
145    z_positions = geometry["z_positions"]
146    tolerance = np.diff(z_positions).min() / 2 if len(z_positions) > 1 else 1.0
147
148    frames = seg.pixel_array
149    if frames.ndim == 2:  # A segmentation with a single frame.
150        frames = frames[None]
151
152    masks = {key: np.zeros(shape, dtype="bool") for key in segment_keys.values()}
153    for frame, frame_group in zip(frames, seg.PerFrameFunctionalGroupsSequence):
154        segment_number = int(frame_group.SegmentIdentificationSequence[0].ReferencedSegmentNumber)
155        key = segment_keys.get(segment_number)
156        if key is None:  # Not one of the 'GTV-1<vis|auto>-<1-5>' segments, e.g. a PET-SUV-threshold-based ROI.
157            continue
158        frame_mask = frame.astype("bool")
159
160        z = None
161        derivation = frame_group.get("DerivationImageSequence", [])
162        if derivation and derivation[0].get("SourceImageSequence"):
163            z = geometry["sop_uids"].get(str(derivation[0].SourceImageSequence[0].ReferencedSOPInstanceUID))
164        if z is None:
165            frame_z = float(frame_group.PlanePositionSequence[0].ImagePositionPatient[2])
166            z = int(np.argmin(np.abs(z_positions - frame_z)))
167            if abs(z_positions[z] - frame_z) > tolerance:
168                if frame_mask.any():
169                    warn(f"Skipping a frame at z={frame_z} in {seg_path}, which does not match a CT slice.")
170                continue
171
172        if flip_axes:
173            frame_mask = np.flip(frame_mask, axis=flip_axes)
174        masks[key][z] |= frame_mask
175
176    return masks
177
178
179def _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir):
180    import h5py
181
182    series_per_subject = defaultdict(dict)
183    for csv_path in csv_paths:
184        with open(csv_path, "r") as f:
185            for row in csv.DictReader(f):
186                series_per_subject[row["Subject ID"]][row["Modality"]] = os.path.join(dicom_dir, row["Series UID"])
187
188    os.makedirs(preprocessed_dir, exist_ok=True)
189    subjects_with_seg = {sid: series for sid, series in series_per_subject.items() if "SEG" in series}
190    for subject_id, series_dirs in tqdm(
191        sorted(subjects_with_seg.items()), desc="Preprocess NSCLC-Radiomics-Interobserver1"
192    ):
193        out_path = os.path.join(preprocessed_dir, f"{subject_id}.h5")
194        if os.path.exists(out_path):
195            continue
196        if "CT" not in series_dirs:
197            warn(f"Skipping {subject_id}, which has a SEG object but no CT series.")
198            continue
199
200        seg_path = glob(os.path.join(series_dirs["SEG"], "*.dcm"))[0]
201        volume, geometry = _load_dicom_volume(series_dirs["CT"], _get_referenced_sop_uids(seg_path))
202        masks = _load_dicom_seg(seg_path, volume.shape, geometry)
203
204        with h5py.File(out_path, "w") as f:
205            f.create_dataset("raw", data=volume, compression="gzip")
206            for key, mask in masks.items():
207                labels = (mask * LABEL_IDS["tumor"]).astype("uint8")
208                f.create_dataset(f"labels/{key}", data=labels, compression="gzip")
209
210
211def get_nsclc_radiomics_interobserver1_data(path: Union[os.PathLike, str], download: bool = False) -> str:
212    """Download the NSCLC-Radiomics-Interobserver1 dataset.
213
214    Args:
215        path: Filepath to a folder where the data is downloaded for further processing.
216        download: Whether to download the data if it is not present.
217
218    Returns:
219        Filepath where the preprocessed data is stored.
220    """
221    # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes.
222    preprocessed_dir = os.path.join(path, "preprocessed")
223    os.makedirs(path, exist_ok=True)
224
225    # Only the segmented series and their referenced CT series are downloaded (see the module docstring),
226    # not the full collection manifest. Each download step is skipped once its metadata csv is written.
227    dicom_dir = os.path.join(path, "dicom")
228    seg_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_seg_series")
229    if not os.path.exists(f"{seg_csv_path}.csv"):
230        if not download:
231            raise RuntimeError(f"Cannot find the data at {dicom_dir}, but download was set to False.")
232        seg_uids = _get_segmented_series_uids()
233        util.download_tcia_series(seg_uids, dicom_dir, seg_csv_path)
234
235    ct_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_ct_series")
236    if not os.path.exists(f"{ct_csv_path}.csv"):
237        with open(f"{seg_csv_path}.csv", "r") as f:
238            seg_uids = [row["Series UID"] for row in csv.DictReader(f)]
239        ct_uids = sorted({
240            _get_referenced_ct_series_uid(glob(os.path.join(dicom_dir, uid, "*.dcm"))[0]) for uid in seg_uids
241        })
242        util.download_tcia_series(ct_uids, dicom_dir, ct_csv_path)
243
244    csv_paths = [f"{seg_csv_path}.csv", f"{ct_csv_path}.csv"]
245    _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir)
246    return preprocessed_dir
247
248
249def get_nsclc_radiomics_interobserver1_paths(
250    path: Union[os.PathLike, str],
251    annotator: int = 1,
252    annotation_type: Literal["manual", "auto"] = "manual",
253    download: bool = False,
254) -> Tuple[List[str], str]:
255    """Get paths to the NSCLC-Radiomics-Interobserver1 data.
256
257    Args:
258        path: Filepath to a folder where the data is downloaded for further processing.
259        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
260            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
261        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
262            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
263            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
264        download: Whether to download the data if it is not present.
265
266    Returns:
267        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data.
268        The key of the label data in each hdf5 file.
269    """
270    assert annotator in ANNOTATORS, f"Invalid annotator: {annotator}. Choose from {ANNOTATORS}."
271    assert annotation_type in _ANNOTATION_TYPE_TO_TAG, f"Invalid annotation_type: {annotation_type}."
272
273    import h5py
274
275    data_dir = get_nsclc_radiomics_interobserver1_data(path, download)
276    label_key = f"labels/{annotation_type}/{annotator}"
277
278    volume_paths = []
279    for volume_path in natsorted(glob(os.path.join(data_dir, "*.h5"))):
280        with h5py.File(volume_path, "r") as f:
281            has_label = label_key in f
282        if has_label:
283            volume_paths.append(volume_path)
284        else:
285            warn(f"Skipping {volume_path}, which has no '{label_key}' delineation.")
286
287    return volume_paths, label_key
288
289
290def get_nsclc_radiomics_interobserver1_dataset(
291    path: Union[os.PathLike, str],
292    patch_shape: Tuple[int, ...],
293    annotator: int = 1,
294    annotation_type: Literal["manual", "auto"] = "manual",
295    resize_inputs: bool = False,
296    download: bool = False,
297    **kwargs
298) -> Dataset:
299    """Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation.
300
301    Args:
302        path: Filepath to a folder where the data is downloaded for further processing.
303        patch_shape: The patch shape to use for training.
304        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
305            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
306        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
307            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
308            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
309        resize_inputs: Whether to resize inputs to the desired patch shape.
310        download: Whether to download the data if it is not present.
311        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
312
313    Returns:
314        The segmentation dataset.
315    """
316    volume_paths, label_key = get_nsclc_radiomics_interobserver1_paths(path, annotator, annotation_type, download)
317
318    if resize_inputs:
319        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
320        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
321            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
322        )
323
324    return torch_em.default_segmentation_dataset(
325        raw_paths=volume_paths,
326        raw_key="raw",
327        label_paths=volume_paths,
328        label_key=label_key,
329        patch_shape=patch_shape,
330        is_seg_dataset=True,
331        **kwargs
332    )
333
334
335def get_nsclc_radiomics_interobserver1_loader(
336    path: Union[os.PathLike, str],
337    batch_size: int,
338    patch_shape: Tuple[int, ...],
339    annotator: int = 1,
340    annotation_type: Literal["manual", "auto"] = "manual",
341    resize_inputs: bool = False,
342    download: bool = False,
343    **kwargs
344) -> DataLoader:
345    """Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation.
346
347    Args:
348        path: Filepath to a folder where the data is downloaded for further processing.
349        batch_size: The batch size for training.
350        patch_shape: The patch shape to use for training.
351        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
352            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
353        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
354            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
355            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
356        resize_inputs: Whether to resize inputs to the desired patch shape.
357        download: Whether to download the data if it is not present.
358        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
359
360    Returns:
361        The DataLoader.
362    """
363    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
364    dataset = get_nsclc_radiomics_interobserver1_dataset(
365        path, patch_shape, annotator, annotation_type, resize_inputs, download, **ds_kwargs
366    )
367    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
COLLECTION = 'NSCLC-Radiomics-Interobserver1'
LABEL_IDS = {'tumor': 1}
ANNOTATORS = (1, 2, 3, 4, 5)
def get_nsclc_radiomics_interobserver1_data(path: Union[os.PathLike, str], download: bool = False) -> str:
212def get_nsclc_radiomics_interobserver1_data(path: Union[os.PathLike, str], download: bool = False) -> str:
213    """Download the NSCLC-Radiomics-Interobserver1 dataset.
214
215    Args:
216        path: Filepath to a folder where the data is downloaded for further processing.
217        download: Whether to download the data if it is not present.
218
219    Returns:
220        Filepath where the preprocessed data is stored.
221    """
222    # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes.
223    preprocessed_dir = os.path.join(path, "preprocessed")
224    os.makedirs(path, exist_ok=True)
225
226    # Only the segmented series and their referenced CT series are downloaded (see the module docstring),
227    # not the full collection manifest. Each download step is skipped once its metadata csv is written.
228    dicom_dir = os.path.join(path, "dicom")
229    seg_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_seg_series")
230    if not os.path.exists(f"{seg_csv_path}.csv"):
231        if not download:
232            raise RuntimeError(f"Cannot find the data at {dicom_dir}, but download was set to False.")
233        seg_uids = _get_segmented_series_uids()
234        util.download_tcia_series(seg_uids, dicom_dir, seg_csv_path)
235
236    ct_csv_path = os.path.join(path, "nsclc_radiomics_interobserver1_ct_series")
237    if not os.path.exists(f"{ct_csv_path}.csv"):
238        with open(f"{seg_csv_path}.csv", "r") as f:
239            seg_uids = [row["Series UID"] for row in csv.DictReader(f)]
240        ct_uids = sorted({
241            _get_referenced_ct_series_uid(glob(os.path.join(dicom_dir, uid, "*.dcm"))[0]) for uid in seg_uids
242        })
243        util.download_tcia_series(ct_uids, dicom_dir, ct_csv_path)
244
245    csv_paths = [f"{seg_csv_path}.csv", f"{ct_csv_path}.csv"]
246    _preprocess_nsclc_radiomics_interobserver1(dicom_dir, csv_paths, preprocessed_dir)
247    return preprocessed_dir

Download the NSCLC-Radiomics-Interobserver1 dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the preprocessed data is stored.

def get_nsclc_radiomics_interobserver1_paths( path: Union[os.PathLike, str], annotator: int = 1, annotation_type: Literal['manual', 'auto'] = 'manual', download: bool = False) -> Tuple[List[str], str]:
250def get_nsclc_radiomics_interobserver1_paths(
251    path: Union[os.PathLike, str],
252    annotator: int = 1,
253    annotation_type: Literal["manual", "auto"] = "manual",
254    download: bool = False,
255) -> Tuple[List[str], str]:
256    """Get paths to the NSCLC-Radiomics-Interobserver1 data.
257
258    Args:
259        path: Filepath to a folder where the data is downloaded for further processing.
260        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
261            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
262        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
263            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
264            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
265        download: Whether to download the data if it is not present.
266
267    Returns:
268        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data.
269        The key of the label data in each hdf5 file.
270    """
271    assert annotator in ANNOTATORS, f"Invalid annotator: {annotator}. Choose from {ANNOTATORS}."
272    assert annotation_type in _ANNOTATION_TYPE_TO_TAG, f"Invalid annotation_type: {annotation_type}."
273
274    import h5py
275
276    data_dir = get_nsclc_radiomics_interobserver1_data(path, download)
277    label_key = f"labels/{annotation_type}/{annotator}"
278
279    volume_paths = []
280    for volume_path in natsorted(glob(os.path.join(data_dir, "*.h5"))):
281        with h5py.File(volume_path, "r") as f:
282            has_label = label_key in f
283        if has_label:
284            volume_paths.append(volume_path)
285        else:
286            warn(f"Skipping {volume_path}, which has no '{label_key}' delineation.")
287
288    return volume_paths, label_key

Get paths to the NSCLC-Radiomics-Interobserver1 data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
  • annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data. The key of the label data in each hdf5 file.

def get_nsclc_radiomics_interobserver1_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], annotator: int = 1, annotation_type: Literal['manual', 'auto'] = 'manual', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
291def get_nsclc_radiomics_interobserver1_dataset(
292    path: Union[os.PathLike, str],
293    patch_shape: Tuple[int, ...],
294    annotator: int = 1,
295    annotation_type: Literal["manual", "auto"] = "manual",
296    resize_inputs: bool = False,
297    download: bool = False,
298    **kwargs
299) -> Dataset:
300    """Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation.
301
302    Args:
303        path: Filepath to a folder where the data is downloaded for further processing.
304        patch_shape: The patch shape to use for training.
305        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
306            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
307        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
308            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
309            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
310        resize_inputs: Whether to resize inputs to the desired patch shape.
311        download: Whether to download the data if it is not present.
312        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
313
314    Returns:
315        The segmentation dataset.
316    """
317    volume_paths, label_key = get_nsclc_radiomics_interobserver1_paths(path, annotator, annotation_type, download)
318
319    if resize_inputs:
320        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
321        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
322            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
323        )
324
325    return torch_em.default_segmentation_dataset(
326        raw_paths=volume_paths,
327        raw_key="raw",
328        label_paths=volume_paths,
329        label_key=label_key,
330        patch_shape=patch_shape,
331        is_seg_dataset=True,
332        **kwargs
333    )

Get the NSCLC-Radiomics-Interobserver1 dataset for lung tumor segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
  • annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_nsclc_radiomics_interobserver1_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], annotator: int = 1, annotation_type: Literal['manual', 'auto'] = 'manual', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
336def get_nsclc_radiomics_interobserver1_loader(
337    path: Union[os.PathLike, str],
338    batch_size: int,
339    patch_shape: Tuple[int, ...],
340    annotator: int = 1,
341    annotation_type: Literal["manual", "auto"] = "manual",
342    resize_inputs: bool = False,
343    download: bool = False,
344    **kwargs
345) -> DataLoader:
346    """Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation.
347
348    Args:
349        path: Filepath to a folder where the data is downloaded for further processing.
350        batch_size: The batch size for training.
351        patch_shape: The patch shape to use for training.
352        annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1'
353            and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
354        annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or
355            'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of
356            the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
357        resize_inputs: Whether to resize inputs to the desired patch shape.
358        download: Whether to download the data if it is not present.
359        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
360
361    Returns:
362        The DataLoader.
363    """
364    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
365    dataset = get_nsclc_radiomics_interobserver1_dataset(
366        path, patch_shape, annotator, annotation_type, resize_inputs, download, **ds_kwargs
367    )
368    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the NSCLC-Radiomics-Interobserver1 dataloader for lung tumor segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • annotator: The radiation oncologist whose delineation to use, an integer from 1 to 5. Oncologists '1' and '3' were trainees at the time of the study, '2', '4' and '5' were extensively experienced.
  • annotation_type: The type of delineation to use. Either 'manual' for a purely manual delineation, or 'auto' for a delineation assisted by an autosegmentation tool and then manually edited. Only 20 of the 21 patients have an 'auto' delineation; the one missing it is skipped for that choice.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.