torch_em.data.datasets.medical.ut_endomri

The UT-EndoMRI dataset contains annotations for pelvic organ segmentation in multi-sequence MRI of endometriosis patients.

The dataset was collected at two clinical institutions. The 'D1' cohort (51 patients, T2-weighted and T1-weighted fat-suppressed sequences) has uterus, ovary, endometrioma, cyst and cul-de-sac structures independently contoured by up to 3 raters. The 'D2' cohort (82 patients, T1-weighted, T1-weighted fat-suppressed, T2-weighted and T2-weighted fat-suppressed sequences) has uterus, ovary and endometrioma structures contoured by a single rater (no rater argument applies to it). Not every sequence or structure is available for every patient, and a label is not tied to a specific sequence in its filename (it may have been contoured on a different sequence than the one requested); get_ut_endomri_paths only returns pairs where the requested sequence and structure annotation both exist and have matching volume shapes.

The data is located at https://doi.org/10.5281/zenodo.13749613. There is no structured license; the record's user agreement states: "The UT-EndoMRI dataset is available for free use exclusively in non-commercial scientific research."

This dataset is from the publication "A Multi-Modal Pelvic MRI Dataset for Deep Learning-Based Pelvic Organ Segmentation in Endometriosis" (Liang et al., submitted). Please cite it if you use this dataset for your research.

  1"""The UT-EndoMRI dataset contains annotations for pelvic organ segmentation in multi-sequence
  2MRI of endometriosis patients.
  3
  4The dataset was collected at two clinical institutions. The 'D1' cohort (51 patients, T2-weighted
  5and T1-weighted fat-suppressed sequences) has uterus, ovary, endometrioma, cyst and cul-de-sac
  6structures independently contoured by up to 3 raters. The 'D2' cohort (82 patients, T1-weighted,
  7T1-weighted fat-suppressed, T2-weighted and T2-weighted fat-suppressed sequences) has uterus,
  8ovary and endometrioma structures contoured by a single rater (no rater argument applies to it).
  9Not every sequence or structure is available for every patient, and a label is not tied to a
 10specific sequence in its filename (it may have been contoured on a different sequence than the
 11one requested); `get_ut_endomri_paths` only returns pairs where the requested sequence and
 12structure annotation both exist and have matching volume shapes.
 13
 14The data is located at https://doi.org/10.5281/zenodo.13749613. There is no structured license;
 15the record's user agreement states: "The UT-EndoMRI dataset is available for free use exclusively
 16in non-commercial scientific research."
 17
 18This dataset is from the publication "A Multi-Modal Pelvic MRI Dataset for Deep Learning-Based
 19Pelvic Organ Segmentation in Endometriosis" (Liang et al., submitted).
 20Please cite it if you use this dataset for your research.
 21"""
 22
 23import os
 24from glob import glob
 25from natsort import natsorted
 26from typing import Union, Tuple, Literal, List, Optional
 27
 28from torch.utils.data import Dataset, DataLoader
 29
 30import torch_em
 31
 32from .. import util
 33
 34
 35URL = "https://zenodo.org/records/13749613/files/UT-EndoMRI.zip"
 36CHECKSUM = "7ab4f9d758c5a2692d78ddf19c35c20b58d77f670a1ec3abb782d6969e278096"
 37
 38DATASETS = {"D1": "D1_MHS", "D2": "D2_TCPW"}
 39SEQUENCES = ("T1", "T1FS", "T2", "T2FS")
 40STRUCTURES = ("ut", "ov", "em", "cy", "cds")
 41
 42
 43def get_ut_endomri_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 44    """Download the UT-EndoMRI dataset.
 45
 46    Args:
 47        path: Filepath to a folder where the data is downloaded for further processing.
 48        download: Whether to download the data if it is not present.
 49
 50    Returns:
 51        Filepath where the data is downloaded.
 52    """
 53    data_dir = os.path.join(path, "UT-EndoMRI")
 54    if os.path.exists(data_dir):
 55        return data_dir
 56
 57    os.makedirs(path, exist_ok=True)
 58
 59    zip_path = os.path.join(path, "UT-EndoMRI.zip")
 60    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
 61    util.unzip(zip_path=zip_path, dst=path)
 62
 63    return data_dir
 64
 65
 66def get_ut_endomri_paths(
 67    path: Union[os.PathLike, str],
 68    dataset: Literal["D1", "D2"] = "D2",
 69    sequence: str = "T2",
 70    structure: str = "ut",
 71    rater: Optional[Literal[1, 2, 3]] = 3,
 72    download: bool = False,
 73) -> Tuple[List[str], List[str]]:
 74    """Get paths to the UT-EndoMRI data.
 75
 76    Args:
 77        path: Filepath to a folder where the data is downloaded for further processing.
 78        dataset: The choice of cohort. Either 'D1' (Memorial Hermann Hospital System) or
 79            'D2' (Texas Children's Hospital Pavilion for Women).
 80        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
 81        structure: The choice of anatomical structure. One of 'ut' (uterus), 'ov' (ovary),
 82            'em' (endometrioma), 'cy' (cyst), 'cds' (cul-de-sac). Not every structure is
 83            annotated for every patient of either cohort.
 84        rater: The choice of annotator for the 'D1' cohort. One of 1, 2 or 3. Ignored for
 85            the 'D2' cohort, which has a single rater.
 86        download: Whether to download the data if it is not present.
 87
 88    Returns:
 89        List of filepaths for the image data.
 90        List of filepaths for the label data.
 91    """
 92    if dataset not in DATASETS:
 93        raise ValueError(f"'{dataset}' is not a valid dataset. Choose one of {list(DATASETS)}.")
 94    if sequence not in SEQUENCES:
 95        raise ValueError(f"'{sequence}' is not a valid sequence. Choose one of {SEQUENCES}.")
 96    if structure not in STRUCTURES:
 97        raise ValueError(f"'{structure}' is not a valid structure. Choose one of {STRUCTURES}.")
 98
 99    import nibabel as nib
100
101    data_dir = get_ut_endomri_data(path, download)
102    cohort_dir = os.path.join(data_dir, DATASETS[dataset])
103
104    label_suffix = f"_{structure}_r{rater}.nii.gz" if dataset == "D1" else f"_{structure}.nii.gz"
105    label_paths = natsorted(glob(os.path.join(cohort_dir, "*", f"*{label_suffix}")))
106
107    raw_paths = []
108    matched_label_paths = []
109    for label_path in label_paths:
110        patient_dir = os.path.dirname(label_path)
111        patient_id = os.path.basename(label_path)[: -len(label_suffix)]
112        raw_path = os.path.join(patient_dir, f"{patient_id}_{sequence}.nii.gz")
113        if not os.path.exists(raw_path):
114            continue
115        # A label is not tied to a specific sequence in its filename, so it may have been
116        # contoured on a different sequence than the one requested here; skip such mismatches.
117        if nib.load(raw_path).shape != nib.load(label_path).shape:
118            continue
119        raw_paths.append(raw_path)
120        matched_label_paths.append(label_path)
121
122    assert len(raw_paths) == len(matched_label_paths) and len(raw_paths) > 0
123    return raw_paths, matched_label_paths
124
125
126def get_ut_endomri_dataset(
127    path: Union[os.PathLike, str],
128    patch_shape: Tuple[int, ...],
129    dataset: Literal["D1", "D2"] = "D2",
130    sequence: str = "T2",
131    structure: str = "ut",
132    rater: Optional[Literal[1, 2, 3]] = 3,
133    resize_inputs: bool = False,
134    download: bool = False,
135    **kwargs
136) -> Dataset:
137    """Get the UT-EndoMRI dataset for pelvic organ segmentation in endometriosis MRI.
138
139    Args:
140        path: Filepath to a folder where the data is downloaded for further processing.
141        patch_shape: The patch shape to use for training.
142        dataset: The choice of cohort. Either 'D1' or 'D2'.
143        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
144        structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
145        rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
146        resize_inputs: Whether to resize the inputs to the patch shape.
147        download: Whether to download the data if it is not present.
148        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
149
150    Returns:
151        The segmentation dataset.
152    """
153    raw_paths, label_paths = get_ut_endomri_paths(path, dataset, sequence, structure, rater, download)
154
155    if resize_inputs:
156        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
157        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
158            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
159        )
160
161    return torch_em.default_segmentation_dataset(
162        raw_paths=raw_paths,
163        raw_key="data",
164        label_paths=label_paths,
165        label_key="data",
166        patch_shape=patch_shape,
167        is_seg_dataset=True,
168        **kwargs
169    )
170
171
172def get_ut_endomri_loader(
173    path: Union[os.PathLike, str],
174    batch_size: int,
175    patch_shape: Tuple[int, ...],
176    dataset: Literal["D1", "D2"] = "D2",
177    sequence: str = "T2",
178    structure: str = "ut",
179    rater: Optional[Literal[1, 2, 3]] = 3,
180    resize_inputs: bool = False,
181    download: bool = False,
182    **kwargs
183) -> DataLoader:
184    """Get the UT-EndoMRI dataloader for pelvic organ segmentation in endometriosis MRI.
185
186    Args:
187        path: Filepath to a folder where the data is downloaded for further processing.
188        batch_size: The batch size for training.
189        patch_shape: The patch shape to use for training.
190        dataset: The choice of cohort. Either 'D1' or 'D2'.
191        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
192        structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
193        rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
194        resize_inputs: Whether to resize the inputs to the patch shape.
195        download: Whether to download the data if it is not present.
196        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
197
198    Returns:
199        The DataLoader.
200    """
201    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
202    dataset_ = get_ut_endomri_dataset(
203        path, patch_shape, dataset, sequence, structure, rater, resize_inputs, download, **ds_kwargs
204    )
205    return torch_em.get_data_loader(dataset_, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/records/13749613/files/UT-EndoMRI.zip'
CHECKSUM = '7ab4f9d758c5a2692d78ddf19c35c20b58d77f670a1ec3abb782d6969e278096'
DATASETS = {'D1': 'D1_MHS', 'D2': 'D2_TCPW'}
SEQUENCES = ('T1', 'T1FS', 'T2', 'T2FS')
STRUCTURES = ('ut', 'ov', 'em', 'cy', 'cds')
def get_ut_endomri_data(path: Union[os.PathLike, str], download: bool = False) -> str:
44def get_ut_endomri_data(path: Union[os.PathLike, str], download: bool = False) -> str:
45    """Download the UT-EndoMRI dataset.
46
47    Args:
48        path: Filepath to a folder where the data is downloaded for further processing.
49        download: Whether to download the data if it is not present.
50
51    Returns:
52        Filepath where the data is downloaded.
53    """
54    data_dir = os.path.join(path, "UT-EndoMRI")
55    if os.path.exists(data_dir):
56        return data_dir
57
58    os.makedirs(path, exist_ok=True)
59
60    zip_path = os.path.join(path, "UT-EndoMRI.zip")
61    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
62    util.unzip(zip_path=zip_path, dst=path)
63
64    return data_dir

Download the UT-EndoMRI dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_ut_endomri_paths( path: Union[os.PathLike, str], dataset: Literal['D1', 'D2'] = 'D2', sequence: str = 'T2', structure: str = 'ut', rater: Optional[Literal[1, 2, 3]] = 3, download: bool = False) -> Tuple[List[str], List[str]]:
 67def get_ut_endomri_paths(
 68    path: Union[os.PathLike, str],
 69    dataset: Literal["D1", "D2"] = "D2",
 70    sequence: str = "T2",
 71    structure: str = "ut",
 72    rater: Optional[Literal[1, 2, 3]] = 3,
 73    download: bool = False,
 74) -> Tuple[List[str], List[str]]:
 75    """Get paths to the UT-EndoMRI data.
 76
 77    Args:
 78        path: Filepath to a folder where the data is downloaded for further processing.
 79        dataset: The choice of cohort. Either 'D1' (Memorial Hermann Hospital System) or
 80            'D2' (Texas Children's Hospital Pavilion for Women).
 81        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
 82        structure: The choice of anatomical structure. One of 'ut' (uterus), 'ov' (ovary),
 83            'em' (endometrioma), 'cy' (cyst), 'cds' (cul-de-sac). Not every structure is
 84            annotated for every patient of either cohort.
 85        rater: The choice of annotator for the 'D1' cohort. One of 1, 2 or 3. Ignored for
 86            the 'D2' cohort, which has a single rater.
 87        download: Whether to download the data if it is not present.
 88
 89    Returns:
 90        List of filepaths for the image data.
 91        List of filepaths for the label data.
 92    """
 93    if dataset not in DATASETS:
 94        raise ValueError(f"'{dataset}' is not a valid dataset. Choose one of {list(DATASETS)}.")
 95    if sequence not in SEQUENCES:
 96        raise ValueError(f"'{sequence}' is not a valid sequence. Choose one of {SEQUENCES}.")
 97    if structure not in STRUCTURES:
 98        raise ValueError(f"'{structure}' is not a valid structure. Choose one of {STRUCTURES}.")
 99
100    import nibabel as nib
101
102    data_dir = get_ut_endomri_data(path, download)
103    cohort_dir = os.path.join(data_dir, DATASETS[dataset])
104
105    label_suffix = f"_{structure}_r{rater}.nii.gz" if dataset == "D1" else f"_{structure}.nii.gz"
106    label_paths = natsorted(glob(os.path.join(cohort_dir, "*", f"*{label_suffix}")))
107
108    raw_paths = []
109    matched_label_paths = []
110    for label_path in label_paths:
111        patient_dir = os.path.dirname(label_path)
112        patient_id = os.path.basename(label_path)[: -len(label_suffix)]
113        raw_path = os.path.join(patient_dir, f"{patient_id}_{sequence}.nii.gz")
114        if not os.path.exists(raw_path):
115            continue
116        # A label is not tied to a specific sequence in its filename, so it may have been
117        # contoured on a different sequence than the one requested here; skip such mismatches.
118        if nib.load(raw_path).shape != nib.load(label_path).shape:
119            continue
120        raw_paths.append(raw_path)
121        matched_label_paths.append(label_path)
122
123    assert len(raw_paths) == len(matched_label_paths) and len(raw_paths) > 0
124    return raw_paths, matched_label_paths

Get paths to the UT-EndoMRI data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • dataset: The choice of cohort. Either 'D1' (Memorial Hermann Hospital System) or 'D2' (Texas Children's Hospital Pavilion for Women).
  • sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
  • structure: The choice of anatomical structure. One of 'ut' (uterus), 'ov' (ovary), 'em' (endometrioma), 'cy' (cyst), 'cds' (cul-de-sac). Not every structure is annotated for every patient of either cohort.
  • rater: The choice of annotator for the 'D1' cohort. One of 1, 2 or 3. Ignored for the 'D2' cohort, which has a single rater.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_ut_endomri_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], dataset: Literal['D1', 'D2'] = 'D2', sequence: str = 'T2', structure: str = 'ut', rater: Optional[Literal[1, 2, 3]] = 3, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
127def get_ut_endomri_dataset(
128    path: Union[os.PathLike, str],
129    patch_shape: Tuple[int, ...],
130    dataset: Literal["D1", "D2"] = "D2",
131    sequence: str = "T2",
132    structure: str = "ut",
133    rater: Optional[Literal[1, 2, 3]] = 3,
134    resize_inputs: bool = False,
135    download: bool = False,
136    **kwargs
137) -> Dataset:
138    """Get the UT-EndoMRI dataset for pelvic organ segmentation in endometriosis MRI.
139
140    Args:
141        path: Filepath to a folder where the data is downloaded for further processing.
142        patch_shape: The patch shape to use for training.
143        dataset: The choice of cohort. Either 'D1' or 'D2'.
144        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
145        structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
146        rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
147        resize_inputs: Whether to resize the inputs to the patch shape.
148        download: Whether to download the data if it is not present.
149        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
150
151    Returns:
152        The segmentation dataset.
153    """
154    raw_paths, label_paths = get_ut_endomri_paths(path, dataset, sequence, structure, rater, download)
155
156    if resize_inputs:
157        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
158        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
159            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
160        )
161
162    return torch_em.default_segmentation_dataset(
163        raw_paths=raw_paths,
164        raw_key="data",
165        label_paths=label_paths,
166        label_key="data",
167        patch_shape=patch_shape,
168        is_seg_dataset=True,
169        **kwargs
170    )

Get the UT-EndoMRI dataset for pelvic organ segmentation in endometriosis MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • dataset: The choice of cohort. Either 'D1' or 'D2'.
  • sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
  • structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
  • rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_ut_endomri_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], dataset: Literal['D1', 'D2'] = 'D2', sequence: str = 'T2', structure: str = 'ut', rater: Optional[Literal[1, 2, 3]] = 3, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
173def get_ut_endomri_loader(
174    path: Union[os.PathLike, str],
175    batch_size: int,
176    patch_shape: Tuple[int, ...],
177    dataset: Literal["D1", "D2"] = "D2",
178    sequence: str = "T2",
179    structure: str = "ut",
180    rater: Optional[Literal[1, 2, 3]] = 3,
181    resize_inputs: bool = False,
182    download: bool = False,
183    **kwargs
184) -> DataLoader:
185    """Get the UT-EndoMRI dataloader for pelvic organ segmentation in endometriosis MRI.
186
187    Args:
188        path: Filepath to a folder where the data is downloaded for further processing.
189        batch_size: The batch size for training.
190        patch_shape: The patch shape to use for training.
191        dataset: The choice of cohort. Either 'D1' or 'D2'.
192        sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
193        structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
194        rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
195        resize_inputs: Whether to resize the inputs to the patch shape.
196        download: Whether to download the data if it is not present.
197        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
198
199    Returns:
200        The DataLoader.
201    """
202    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
203    dataset_ = get_ut_endomri_dataset(
204        path, patch_shape, dataset, sequence, structure, rater, resize_inputs, download, **ds_kwargs
205    )
206    return torch_em.get_data_loader(dataset_, batch_size, **loader_kwargs)

Get the UT-EndoMRI dataloader for pelvic organ segmentation in endometriosis MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • dataset: The choice of cohort. Either 'D1' or 'D2'.
  • sequence: The choice of MRI sequence. One of 'T1', 'T1FS', 'T2', 'T2FS'.
  • structure: The choice of anatomical structure. One of 'ut', 'ov', 'em', 'cy', 'cds'.
  • rater: The choice of annotator for the 'D1' cohort. Ignored for the 'D2' cohort.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.