torch_em.data.datasets.medical.retouch

The RETOUCH dataset contains annotations for intraretinal fluid (IRF), subretinal fluid (SRF) and pigment epithelial detachment (PED) segmentation in 3D retinal OCT volumes.

The dataset was curated for the RETOUCH challenge (MICCAI 2017, https://retouch.grand-challenge.org). It comprises 112 OCT volumes acquired with three different device manufacturers (Cirrus, Spectralis and Topcon), out of which 70 volumes (24 Cirrus, 24 Spectralis, 22 Topcon) form the labeled training set and 42 volumes form the unlabeled test set. Only the labeled training set is exposed here, as the test set annotations are not publicly available.

NOTE: The dataset is gated and cannot be downloaded automatically by 'torch_em'. To obtain it:

  • Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.
  • Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' document.
  • Once your registration is approved (contact hrvoje.bogunovic@meduniwien.ac.at for questions), you will get access to the 'Download' page, from which the training data must be downloaded.
  • Extract the downloaded archive(s) to the desired 'path'. The expected layout per vendor is: '/RETOUCH-TrainingSet-/TRAIN/oct.mhd', 'oct.raw', 'reference.mhd', 'reference.raw', where '' is one of 'Cirrus', 'Spectralis', 'Topcon'.

The original MetaImage ('.mhd' / '.raw') volumes are converted to hdf5 once and then loaded from there.

The reference annotations use the label ids: 0 = background, 1 = IRF, 2 = SRF, 3 = PED.

The dataset is from the publication https://doi.org/10.1109/TMI.2019.2901398. Please cite it if you use this dataset in your research.

  1"""The RETOUCH dataset contains annotations for intraretinal fluid (IRF), subretinal fluid (SRF)
  2and pigment epithelial detachment (PED) segmentation in 3D retinal OCT volumes.
  3
  4The dataset was curated for the RETOUCH challenge (MICCAI 2017, https://retouch.grand-challenge.org).
  5It comprises 112 OCT volumes acquired with three different device manufacturers (Cirrus, Spectralis
  6and Topcon), out of which 70 volumes (24 Cirrus, 24 Spectralis, 22 Topcon) form the labeled training
  7set and 42 volumes form the unlabeled test set. Only the labeled training set is exposed here, as the
  8test set annotations are not publicly available.
  9
 10NOTE: The dataset is gated and cannot be downloaded automatically by 'torch_em'. To obtain it:
 11- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.
 12- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' document.
 13- Once your registration is approved (contact hrvoje.bogunovic@meduniwien.ac.at for questions),
 14  you will get access to the 'Download' page, from which the training data must be downloaded.
 15- Extract the downloaded archive(s) to the desired 'path'. The expected layout per vendor is:
 16  '<path>/RETOUCH-TrainingSet-<vendor>/TRAIN<id>/oct.mhd', 'oct.raw', 'reference.mhd', 'reference.raw',
 17  where '<vendor>' is one of 'Cirrus', 'Spectralis', 'Topcon'.
 18
 19The original MetaImage ('.mhd' / '.raw') volumes are converted to hdf5 once and then loaded from there.
 20
 21The reference annotations use the label ids: 0 = background, 1 = IRF, 2 = SRF, 3 = PED.
 22
 23The dataset is from the publication https://doi.org/10.1109/TMI.2019.2901398.
 24Please cite it if you use this dataset in your research.
 25"""
 26
 27import os
 28from glob import glob
 29from tqdm import tqdm
 30from natsort import natsorted
 31from typing import Union, Tuple, List, Literal, Sequence
 32
 33import numpy as np
 34
 35from torch.utils.data import Dataset, DataLoader
 36
 37import torch_em
 38
 39from .. import util
 40
 41
 42VENDORS = ["Cirrus", "Spectralis", "Topcon"]
 43
 44METAIMAGE_DTYPES = {
 45    "MET_CHAR": "int8", "MET_UCHAR": "uint8", "MET_SHORT": "int16", "MET_USHORT": "uint16",
 46    "MET_INT": "int32", "MET_UINT": "uint32", "MET_FLOAT": "float32", "MET_DOUBLE": "float64",
 47}
 48
 49
 50def _read_metaimage(path):
 51    """Read an uncompressed MetaImage ('.mhd' + '.raw') volume with numpy, returning it in 'zyx' order."""
 52    header = {}
 53    with open(path, "r") as f:
 54        for line in f:
 55            if "=" in line:
 56                key, value = line.split("=", 1)
 57                header[key.strip()] = value.strip()
 58
 59    if header.get("CompressedData", "False") == "True":
 60        raise NotImplementedError(f"Compressed MetaImage data is not supported: '{path}'.")
 61
 62    shape = tuple(int(s) for s in header["DimSize"].split())[::-1]
 63    dtype = np.dtype(METAIMAGE_DTYPES[header["ElementType"]])
 64    if header.get("BinaryDataByteOrderMSB", "False") == "True":
 65        dtype = dtype.newbyteorder(">")
 66
 67    raw_path = os.path.join(os.path.dirname(path), header["ElementDataFile"])
 68    return np.fromfile(raw_path, dtype=dtype).reshape(shape)
 69
 70
 71def get_retouch_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 72    """Obtain the RETOUCH dataset.
 73
 74    Args:
 75        path: Filepath to a folder where the data is downloaded for further processing.
 76        download: Whether to download the data if it is not present.
 77
 78    Returns:
 79        Filepath to the folder with the converted hdf5 volumes.
 80    """
 81    data_dir = os.path.join(path, "preprocessed")
 82    if os.path.exists(data_dir):
 83        return data_dir
 84
 85    vendor_dirs = [os.path.join(path, f"RETOUCH-TrainingSet-{vendor}") for vendor in VENDORS]
 86    if not any(os.path.exists(vendor_dir) for vendor_dir in vendor_dirs):
 87        msg = "The RETOUCH dataset is gated and cannot be downloaded automatically. To obtain it:\n" \
 88            "- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.\n" \
 89            "- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' " \
 90            "document. Contact hrvoje.bogunovic@meduniwien.ac.at for registration questions.\n" \
 91            "- Once approved, download the training data from the challenge 'Download' page.\n" \
 92            f"- Extract the downloaded archive(s) so that '{path}' contains one folder per vendor, " \
 93            "named 'RETOUCH-TrainingSet-Cirrus', 'RETOUCH-TrainingSet-Spectralis' and " \
 94            "'RETOUCH-TrainingSet-Topcon', each with 'TRAIN<id>' subfolders containing " \
 95            "'oct.mhd', 'oct.raw', 'reference.mhd' and 'reference.raw'."
 96        if download:
 97            msg = "Download is set to True, but 'torch_em' cannot download this dataset.\n" + msg
 98        raise RuntimeError(msg)
 99
100    import h5py
101
102    os.makedirs(data_dir, exist_ok=True)
103    for vendor_dir in vendor_dirs:
104        if not os.path.exists(vendor_dir):
105            continue
106
107        vendor = os.path.basename(vendor_dir).replace("RETOUCH-TrainingSet-", "")
108        case_dirs = natsorted(glob(os.path.join(vendor_dir, "TRAIN*")))
109        for case_dir in tqdm(case_dirs, desc=f"Converting RETOUCH volumes for '{os.path.basename(vendor_dir)}'"):
110            case_id = os.path.basename(case_dir)
111            out_path = os.path.join(data_dir, f"{vendor}_{case_id}.h5")
112            if os.path.exists(out_path):
113                continue
114
115            raw = _read_metaimage(os.path.join(case_dir, "oct.mhd"))
116            labels = _read_metaimage(os.path.join(case_dir, "reference.mhd")).astype("uint8")
117            assert raw.shape == labels.shape, f"Shape mismatch for '{case_dir}': {raw.shape} vs. {labels.shape}."
118
119            with h5py.File(out_path, "w") as f:
120                f.create_dataset("raw", data=raw, compression="gzip")
121                f.create_dataset("labels", data=labels, compression="gzip")
122
123    return data_dir
124
125
126def get_retouch_paths(
127    path: Union[os.PathLike, str], vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
128    download: bool = False,
129) -> List[str]:
130    """Get paths to the RETOUCH data.
131
132    Args:
133        path: Filepath to a folder where the data is downloaded for further processing.
134        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
135        download: Whether to download the data if it is not present.
136
137    Returns:
138        List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'.
139    """
140    data_dir = get_retouch_data(path, download)
141
142    if vendor is None:
143        vendors = VENDORS
144    else:
145        vendors = [vendor] if isinstance(vendor, str) else list(vendor)
146        for v in vendors:
147            if v not in VENDORS:
148                raise ValueError(f"'{v}' is not a valid vendor. Choose from {VENDORS}.")
149
150    volume_paths = []
151    for v in vendors:
152        volume_paths.extend(natsorted(glob(os.path.join(data_dir, f"{v}_TRAIN*.h5"))))
153
154    assert len(volume_paths) > 0, f"No RETOUCH volumes found at '{data_dir}'."
155    return volume_paths
156
157
158def get_retouch_dataset(
159    path: Union[os.PathLike, str],
160    patch_shape: Tuple[int, ...],
161    vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
162    resize_inputs: bool = False,
163    download: bool = False,
164    **kwargs
165) -> Dataset:
166    """Get the RETOUCH dataset for retinal fluid segmentation.
167
168    Args:
169        path: Filepath to a folder where the data is downloaded for further processing.
170        patch_shape: The patch shape to use for training.
171        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
172        resize_inputs: Whether to resize inputs to the desired patch shape.
173        download: Whether to download the data if it is not present.
174        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
175
176    Returns:
177        The segmentation dataset.
178    """
179    volume_paths = get_retouch_paths(path, vendor, download)
180
181    if resize_inputs:
182        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
183        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
184            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
185        )
186
187    return torch_em.default_segmentation_dataset(
188        raw_paths=volume_paths,
189        raw_key="raw",
190        label_paths=volume_paths,
191        label_key="labels",
192        patch_shape=patch_shape,
193        is_seg_dataset=True,
194        **kwargs
195    )
196
197
198def get_retouch_loader(
199    path: Union[os.PathLike, str],
200    batch_size: int,
201    patch_shape: Tuple[int, ...],
202    vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
203    resize_inputs: bool = False,
204    download: bool = False,
205    **kwargs
206) -> DataLoader:
207    """Get the RETOUCH dataloader for retinal fluid segmentation.
208
209    Args:
210        path: Filepath to a folder where the data is downloaded for further processing.
211        batch_size: The batch size for training.
212        patch_shape: The patch shape to use for training.
213        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
214        resize_inputs: Whether to resize inputs to the desired patch shape.
215        download: Whether to download the data if it is not present.
216        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
217
218    Returns:
219        The DataLoader.
220    """
221    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
222    dataset = get_retouch_dataset(path, patch_shape, vendor, resize_inputs, download, **ds_kwargs)
223    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
VENDORS = ['Cirrus', 'Spectralis', 'Topcon']
METAIMAGE_DTYPES = {'MET_CHAR': 'int8', 'MET_UCHAR': 'uint8', 'MET_SHORT': 'int16', 'MET_USHORT': 'uint16', 'MET_INT': 'int32', 'MET_UINT': 'uint32', 'MET_FLOAT': 'float32', 'MET_DOUBLE': 'float64'}
def get_retouch_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 72def get_retouch_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 73    """Obtain the RETOUCH dataset.
 74
 75    Args:
 76        path: Filepath to a folder where the data is downloaded for further processing.
 77        download: Whether to download the data if it is not present.
 78
 79    Returns:
 80        Filepath to the folder with the converted hdf5 volumes.
 81    """
 82    data_dir = os.path.join(path, "preprocessed")
 83    if os.path.exists(data_dir):
 84        return data_dir
 85
 86    vendor_dirs = [os.path.join(path, f"RETOUCH-TrainingSet-{vendor}") for vendor in VENDORS]
 87    if not any(os.path.exists(vendor_dir) for vendor_dir in vendor_dirs):
 88        msg = "The RETOUCH dataset is gated and cannot be downloaded automatically. To obtain it:\n" \
 89            "- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.\n" \
 90            "- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' " \
 91            "document. Contact hrvoje.bogunovic@meduniwien.ac.at for registration questions.\n" \
 92            "- Once approved, download the training data from the challenge 'Download' page.\n" \
 93            f"- Extract the downloaded archive(s) so that '{path}' contains one folder per vendor, " \
 94            "named 'RETOUCH-TrainingSet-Cirrus', 'RETOUCH-TrainingSet-Spectralis' and " \
 95            "'RETOUCH-TrainingSet-Topcon', each with 'TRAIN<id>' subfolders containing " \
 96            "'oct.mhd', 'oct.raw', 'reference.mhd' and 'reference.raw'."
 97        if download:
 98            msg = "Download is set to True, but 'torch_em' cannot download this dataset.\n" + msg
 99        raise RuntimeError(msg)
100
101    import h5py
102
103    os.makedirs(data_dir, exist_ok=True)
104    for vendor_dir in vendor_dirs:
105        if not os.path.exists(vendor_dir):
106            continue
107
108        vendor = os.path.basename(vendor_dir).replace("RETOUCH-TrainingSet-", "")
109        case_dirs = natsorted(glob(os.path.join(vendor_dir, "TRAIN*")))
110        for case_dir in tqdm(case_dirs, desc=f"Converting RETOUCH volumes for '{os.path.basename(vendor_dir)}'"):
111            case_id = os.path.basename(case_dir)
112            out_path = os.path.join(data_dir, f"{vendor}_{case_id}.h5")
113            if os.path.exists(out_path):
114                continue
115
116            raw = _read_metaimage(os.path.join(case_dir, "oct.mhd"))
117            labels = _read_metaimage(os.path.join(case_dir, "reference.mhd")).astype("uint8")
118            assert raw.shape == labels.shape, f"Shape mismatch for '{case_dir}': {raw.shape} vs. {labels.shape}."
119
120            with h5py.File(out_path, "w") as f:
121                f.create_dataset("raw", data=raw, compression="gzip")
122                f.create_dataset("labels", data=labels, compression="gzip")
123
124    return data_dir

Obtain the RETOUCH dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder with the converted hdf5 volumes.

def get_retouch_paths( path: Union[os.PathLike, str], vendor: Union[Literal['Cirrus', 'Spectralis', 'Topcon'], Sequence[str], NoneType] = None, download: bool = False) -> List[str]:
127def get_retouch_paths(
128    path: Union[os.PathLike, str], vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
129    download: bool = False,
130) -> List[str]:
131    """Get paths to the RETOUCH data.
132
133    Args:
134        path: Filepath to a folder where the data is downloaded for further processing.
135        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
136        download: Whether to download the data if it is not present.
137
138    Returns:
139        List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'.
140    """
141    data_dir = get_retouch_data(path, download)
142
143    if vendor is None:
144        vendors = VENDORS
145    else:
146        vendors = [vendor] if isinstance(vendor, str) else list(vendor)
147        for v in vendors:
148            if v not in VENDORS:
149                raise ValueError(f"'{v}' is not a valid vendor. Choose from {VENDORS}.")
150
151    volume_paths = []
152    for v in vendors:
153        volume_paths.extend(natsorted(glob(os.path.join(data_dir, f"{v}_TRAIN*.h5"))))
154
155    assert len(volume_paths) > 0, f"No RETOUCH volumes found at '{data_dir}'."
156    return volume_paths

Get paths to the RETOUCH data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'.

def get_retouch_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], vendor: Union[Literal['Cirrus', 'Spectralis', 'Topcon'], Sequence[str], NoneType] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
159def get_retouch_dataset(
160    path: Union[os.PathLike, str],
161    patch_shape: Tuple[int, ...],
162    vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
163    resize_inputs: bool = False,
164    download: bool = False,
165    **kwargs
166) -> Dataset:
167    """Get the RETOUCH dataset for retinal fluid segmentation.
168
169    Args:
170        path: Filepath to a folder where the data is downloaded for further processing.
171        patch_shape: The patch shape to use for training.
172        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
173        resize_inputs: Whether to resize inputs to the desired patch shape.
174        download: Whether to download the data if it is not present.
175        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
176
177    Returns:
178        The segmentation dataset.
179    """
180    volume_paths = get_retouch_paths(path, vendor, download)
181
182    if resize_inputs:
183        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
184        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
185            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
186        )
187
188    return torch_em.default_segmentation_dataset(
189        raw_paths=volume_paths,
190        raw_key="raw",
191        label_paths=volume_paths,
192        label_key="labels",
193        patch_shape=patch_shape,
194        is_seg_dataset=True,
195        **kwargs
196    )

Get the RETOUCH dataset for retinal fluid segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_retouch_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], vendor: Union[Literal['Cirrus', 'Spectralis', 'Topcon'], Sequence[str], NoneType] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
199def get_retouch_loader(
200    path: Union[os.PathLike, str],
201    batch_size: int,
202    patch_shape: Tuple[int, ...],
203    vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None,
204    resize_inputs: bool = False,
205    download: bool = False,
206    **kwargs
207) -> DataLoader:
208    """Get the RETOUCH dataloader for retinal fluid segmentation.
209
210    Args:
211        path: Filepath to a folder where the data is downloaded for further processing.
212        batch_size: The batch size for training.
213        patch_shape: The patch shape to use for training.
214        vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
215        resize_inputs: Whether to resize inputs to the desired patch shape.
216        download: Whether to download the data if it is not present.
217        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
218
219    Returns:
220        The DataLoader.
221    """
222    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
223    dataset = get_retouch_dataset(path, patch_shape, vendor, resize_inputs, download, **ds_kwargs)
224    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the RETOUCH dataloader for retinal fluid segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.