torch_em.data.datasets.medical.retouch
The RETOUCH dataset contains annotations for intraretinal fluid (IRF), subretinal fluid (SRF) and pigment epithelial detachment (PED) segmentation in 3D retinal OCT volumes.
The dataset was curated for the RETOUCH challenge (MICCAI 2017, https://retouch.grand-challenge.org). It comprises 112 OCT volumes acquired with three different device manufacturers (Cirrus, Spectralis and Topcon), out of which 70 volumes (24 Cirrus, 24 Spectralis, 22 Topcon) form the labeled training set and 42 volumes form the unlabeled test set. Only the labeled training set is exposed here, as the test set annotations are not publicly available.
NOTE: The dataset is gated and cannot be downloaded automatically by 'torch_em'. To obtain it:
- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.
- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' document.
- Once your registration is approved (contact hrvoje.bogunovic@meduniwien.ac.at for questions), you will get access to the 'Download' page, from which the training data must be downloaded.
- Extract the downloaded archive(s) to the desired 'path'. The expected layout per vendor is:
'
/RETOUCH-TrainingSet- /TRAIN /oct.mhd', 'oct.raw', 'reference.mhd', 'reference.raw', where ' ' is one of 'Cirrus', 'Spectralis', 'Topcon'.
The original MetaImage ('.mhd' / '.raw') volumes are converted to hdf5 once and then loaded from there.
The reference annotations use the label ids: 0 = background, 1 = IRF, 2 = SRF, 3 = PED.
The dataset is from the publication https://doi.org/10.1109/TMI.2019.2901398. Please cite it if you use this dataset in your research.
1"""The RETOUCH dataset contains annotations for intraretinal fluid (IRF), subretinal fluid (SRF) 2and pigment epithelial detachment (PED) segmentation in 3D retinal OCT volumes. 3 4The dataset was curated for the RETOUCH challenge (MICCAI 2017, https://retouch.grand-challenge.org). 5It comprises 112 OCT volumes acquired with three different device manufacturers (Cirrus, Spectralis 6and Topcon), out of which 70 volumes (24 Cirrus, 24 Spectralis, 22 Topcon) form the labeled training 7set and 42 volumes form the unlabeled test set. Only the labeled training set is exposed here, as the 8test set annotations are not publicly available. 9 10NOTE: The dataset is gated and cannot be downloaded automatically by 'torch_em'. To obtain it: 11- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge. 12- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' document. 13- Once your registration is approved (contact hrvoje.bogunovic@meduniwien.ac.at for questions), 14 you will get access to the 'Download' page, from which the training data must be downloaded. 15- Extract the downloaded archive(s) to the desired 'path'. The expected layout per vendor is: 16 '<path>/RETOUCH-TrainingSet-<vendor>/TRAIN<id>/oct.mhd', 'oct.raw', 'reference.mhd', 'reference.raw', 17 where '<vendor>' is one of 'Cirrus', 'Spectralis', 'Topcon'. 18 19The original MetaImage ('.mhd' / '.raw') volumes are converted to hdf5 once and then loaded from there. 20 21The reference annotations use the label ids: 0 = background, 1 = IRF, 2 = SRF, 3 = PED. 22 23The dataset is from the publication https://doi.org/10.1109/TMI.2019.2901398. 24Please cite it if you use this dataset in your research. 25""" 26 27import os 28from glob import glob 29from tqdm import tqdm 30from natsort import natsorted 31from typing import Union, Tuple, List, Literal, Sequence 32 33import numpy as np 34 35from torch.utils.data import Dataset, DataLoader 36 37import torch_em 38 39from .. import util 40 41 42VENDORS = ["Cirrus", "Spectralis", "Topcon"] 43 44METAIMAGE_DTYPES = { 45 "MET_CHAR": "int8", "MET_UCHAR": "uint8", "MET_SHORT": "int16", "MET_USHORT": "uint16", 46 "MET_INT": "int32", "MET_UINT": "uint32", "MET_FLOAT": "float32", "MET_DOUBLE": "float64", 47} 48 49 50def _read_metaimage(path): 51 """Read an uncompressed MetaImage ('.mhd' + '.raw') volume with numpy, returning it in 'zyx' order.""" 52 header = {} 53 with open(path, "r") as f: 54 for line in f: 55 if "=" in line: 56 key, value = line.split("=", 1) 57 header[key.strip()] = value.strip() 58 59 if header.get("CompressedData", "False") == "True": 60 raise NotImplementedError(f"Compressed MetaImage data is not supported: '{path}'.") 61 62 shape = tuple(int(s) for s in header["DimSize"].split())[::-1] 63 dtype = np.dtype(METAIMAGE_DTYPES[header["ElementType"]]) 64 if header.get("BinaryDataByteOrderMSB", "False") == "True": 65 dtype = dtype.newbyteorder(">") 66 67 raw_path = os.path.join(os.path.dirname(path), header["ElementDataFile"]) 68 return np.fromfile(raw_path, dtype=dtype).reshape(shape) 69 70 71def get_retouch_data(path: Union[os.PathLike, str], download: bool = False) -> str: 72 """Obtain the RETOUCH dataset. 73 74 Args: 75 path: Filepath to a folder where the data is downloaded for further processing. 76 download: Whether to download the data if it is not present. 77 78 Returns: 79 Filepath to the folder with the converted hdf5 volumes. 80 """ 81 data_dir = os.path.join(path, "preprocessed") 82 if os.path.exists(data_dir): 83 return data_dir 84 85 vendor_dirs = [os.path.join(path, f"RETOUCH-TrainingSet-{vendor}") for vendor in VENDORS] 86 if not any(os.path.exists(vendor_dir) for vendor_dir in vendor_dirs): 87 msg = "The RETOUCH dataset is gated and cannot be downloaded automatically. To obtain it:\n" \ 88 "- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.\n" \ 89 "- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' " \ 90 "document. Contact hrvoje.bogunovic@meduniwien.ac.at for registration questions.\n" \ 91 "- Once approved, download the training data from the challenge 'Download' page.\n" \ 92 f"- Extract the downloaded archive(s) so that '{path}' contains one folder per vendor, " \ 93 "named 'RETOUCH-TrainingSet-Cirrus', 'RETOUCH-TrainingSet-Spectralis' and " \ 94 "'RETOUCH-TrainingSet-Topcon', each with 'TRAIN<id>' subfolders containing " \ 95 "'oct.mhd', 'oct.raw', 'reference.mhd' and 'reference.raw'." 96 if download: 97 msg = "Download is set to True, but 'torch_em' cannot download this dataset.\n" + msg 98 raise RuntimeError(msg) 99 100 import h5py 101 102 os.makedirs(data_dir, exist_ok=True) 103 for vendor_dir in vendor_dirs: 104 if not os.path.exists(vendor_dir): 105 continue 106 107 vendor = os.path.basename(vendor_dir).replace("RETOUCH-TrainingSet-", "") 108 case_dirs = natsorted(glob(os.path.join(vendor_dir, "TRAIN*"))) 109 for case_dir in tqdm(case_dirs, desc=f"Converting RETOUCH volumes for '{os.path.basename(vendor_dir)}'"): 110 case_id = os.path.basename(case_dir) 111 out_path = os.path.join(data_dir, f"{vendor}_{case_id}.h5") 112 if os.path.exists(out_path): 113 continue 114 115 raw = _read_metaimage(os.path.join(case_dir, "oct.mhd")) 116 labels = _read_metaimage(os.path.join(case_dir, "reference.mhd")).astype("uint8") 117 assert raw.shape == labels.shape, f"Shape mismatch for '{case_dir}': {raw.shape} vs. {labels.shape}." 118 119 with h5py.File(out_path, "w") as f: 120 f.create_dataset("raw", data=raw, compression="gzip") 121 f.create_dataset("labels", data=labels, compression="gzip") 122 123 return data_dir 124 125 126def get_retouch_paths( 127 path: Union[os.PathLike, str], vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 128 download: bool = False, 129) -> List[str]: 130 """Get paths to the RETOUCH data. 131 132 Args: 133 path: Filepath to a folder where the data is downloaded for further processing. 134 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 135 download: Whether to download the data if it is not present. 136 137 Returns: 138 List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'. 139 """ 140 data_dir = get_retouch_data(path, download) 141 142 if vendor is None: 143 vendors = VENDORS 144 else: 145 vendors = [vendor] if isinstance(vendor, str) else list(vendor) 146 for v in vendors: 147 if v not in VENDORS: 148 raise ValueError(f"'{v}' is not a valid vendor. Choose from {VENDORS}.") 149 150 volume_paths = [] 151 for v in vendors: 152 volume_paths.extend(natsorted(glob(os.path.join(data_dir, f"{v}_TRAIN*.h5")))) 153 154 assert len(volume_paths) > 0, f"No RETOUCH volumes found at '{data_dir}'." 155 return volume_paths 156 157 158def get_retouch_dataset( 159 path: Union[os.PathLike, str], 160 patch_shape: Tuple[int, ...], 161 vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 162 resize_inputs: bool = False, 163 download: bool = False, 164 **kwargs 165) -> Dataset: 166 """Get the RETOUCH dataset for retinal fluid segmentation. 167 168 Args: 169 path: Filepath to a folder where the data is downloaded for further processing. 170 patch_shape: The patch shape to use for training. 171 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 172 resize_inputs: Whether to resize inputs to the desired patch shape. 173 download: Whether to download the data if it is not present. 174 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 175 176 Returns: 177 The segmentation dataset. 178 """ 179 volume_paths = get_retouch_paths(path, vendor, download) 180 181 if resize_inputs: 182 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 183 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 184 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 185 ) 186 187 return torch_em.default_segmentation_dataset( 188 raw_paths=volume_paths, 189 raw_key="raw", 190 label_paths=volume_paths, 191 label_key="labels", 192 patch_shape=patch_shape, 193 is_seg_dataset=True, 194 **kwargs 195 ) 196 197 198def get_retouch_loader( 199 path: Union[os.PathLike, str], 200 batch_size: int, 201 patch_shape: Tuple[int, ...], 202 vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 203 resize_inputs: bool = False, 204 download: bool = False, 205 **kwargs 206) -> DataLoader: 207 """Get the RETOUCH dataloader for retinal fluid segmentation. 208 209 Args: 210 path: Filepath to a folder where the data is downloaded for further processing. 211 batch_size: The batch size for training. 212 patch_shape: The patch shape to use for training. 213 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 214 resize_inputs: Whether to resize inputs to the desired patch shape. 215 download: Whether to download the data if it is not present. 216 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 217 218 Returns: 219 The DataLoader. 220 """ 221 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 222 dataset = get_retouch_dataset(path, patch_shape, vendor, resize_inputs, download, **ds_kwargs) 223 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
72def get_retouch_data(path: Union[os.PathLike, str], download: bool = False) -> str: 73 """Obtain the RETOUCH dataset. 74 75 Args: 76 path: Filepath to a folder where the data is downloaded for further processing. 77 download: Whether to download the data if it is not present. 78 79 Returns: 80 Filepath to the folder with the converted hdf5 volumes. 81 """ 82 data_dir = os.path.join(path, "preprocessed") 83 if os.path.exists(data_dir): 84 return data_dir 85 86 vendor_dirs = [os.path.join(path, f"RETOUCH-TrainingSet-{vendor}") for vendor in VENDORS] 87 if not any(os.path.exists(vendor_dir) for vendor_dir in vendor_dirs): 88 msg = "The RETOUCH dataset is gated and cannot be downloaded automatically. To obtain it:\n" \ 89 "- Visit https://retouch.grand-challenge.org and click on 'Join' to register for the challenge.\n" \ 90 "- Complete the registration form and sign the 'RETOUCH-Agreement_of_Data_Confidentiality' " \ 91 "document. Contact hrvoje.bogunovic@meduniwien.ac.at for registration questions.\n" \ 92 "- Once approved, download the training data from the challenge 'Download' page.\n" \ 93 f"- Extract the downloaded archive(s) so that '{path}' contains one folder per vendor, " \ 94 "named 'RETOUCH-TrainingSet-Cirrus', 'RETOUCH-TrainingSet-Spectralis' and " \ 95 "'RETOUCH-TrainingSet-Topcon', each with 'TRAIN<id>' subfolders containing " \ 96 "'oct.mhd', 'oct.raw', 'reference.mhd' and 'reference.raw'." 97 if download: 98 msg = "Download is set to True, but 'torch_em' cannot download this dataset.\n" + msg 99 raise RuntimeError(msg) 100 101 import h5py 102 103 os.makedirs(data_dir, exist_ok=True) 104 for vendor_dir in vendor_dirs: 105 if not os.path.exists(vendor_dir): 106 continue 107 108 vendor = os.path.basename(vendor_dir).replace("RETOUCH-TrainingSet-", "") 109 case_dirs = natsorted(glob(os.path.join(vendor_dir, "TRAIN*"))) 110 for case_dir in tqdm(case_dirs, desc=f"Converting RETOUCH volumes for '{os.path.basename(vendor_dir)}'"): 111 case_id = os.path.basename(case_dir) 112 out_path = os.path.join(data_dir, f"{vendor}_{case_id}.h5") 113 if os.path.exists(out_path): 114 continue 115 116 raw = _read_metaimage(os.path.join(case_dir, "oct.mhd")) 117 labels = _read_metaimage(os.path.join(case_dir, "reference.mhd")).astype("uint8") 118 assert raw.shape == labels.shape, f"Shape mismatch for '{case_dir}': {raw.shape} vs. {labels.shape}." 119 120 with h5py.File(out_path, "w") as f: 121 f.create_dataset("raw", data=raw, compression="gzip") 122 f.create_dataset("labels", data=labels, compression="gzip") 123 124 return data_dir
Obtain the RETOUCH dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder with the converted hdf5 volumes.
127def get_retouch_paths( 128 path: Union[os.PathLike, str], vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 129 download: bool = False, 130) -> List[str]: 131 """Get paths to the RETOUCH data. 132 133 Args: 134 path: Filepath to a folder where the data is downloaded for further processing. 135 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 136 download: Whether to download the data if it is not present. 137 138 Returns: 139 List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'. 140 """ 141 data_dir = get_retouch_data(path, download) 142 143 if vendor is None: 144 vendors = VENDORS 145 else: 146 vendors = [vendor] if isinstance(vendor, str) else list(vendor) 147 for v in vendors: 148 if v not in VENDORS: 149 raise ValueError(f"'{v}' is not a valid vendor. Choose from {VENDORS}.") 150 151 volume_paths = [] 152 for v in vendors: 153 volume_paths.extend(natsorted(glob(os.path.join(data_dir, f"{v}_TRAIN*.h5")))) 154 155 assert len(volume_paths) > 0, f"No RETOUCH volumes found at '{data_dir}'." 156 return volume_paths
Get paths to the RETOUCH data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 volumes, which store the image data at 'raw' and the label data at 'labels'.
159def get_retouch_dataset( 160 path: Union[os.PathLike, str], 161 patch_shape: Tuple[int, ...], 162 vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 163 resize_inputs: bool = False, 164 download: bool = False, 165 **kwargs 166) -> Dataset: 167 """Get the RETOUCH dataset for retinal fluid segmentation. 168 169 Args: 170 path: Filepath to a folder where the data is downloaded for further processing. 171 patch_shape: The patch shape to use for training. 172 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 173 resize_inputs: Whether to resize inputs to the desired patch shape. 174 download: Whether to download the data if it is not present. 175 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 176 177 Returns: 178 The segmentation dataset. 179 """ 180 volume_paths = get_retouch_paths(path, vendor, download) 181 182 if resize_inputs: 183 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 184 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 185 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 186 ) 187 188 return torch_em.default_segmentation_dataset( 189 raw_paths=volume_paths, 190 raw_key="raw", 191 label_paths=volume_paths, 192 label_key="labels", 193 patch_shape=patch_shape, 194 is_seg_dataset=True, 195 **kwargs 196 )
Get the RETOUCH dataset for retinal fluid segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
199def get_retouch_loader( 200 path: Union[os.PathLike, str], 201 batch_size: int, 202 patch_shape: Tuple[int, ...], 203 vendor: Union[Literal["Cirrus", "Spectralis", "Topcon"], Sequence[str], None] = None, 204 resize_inputs: bool = False, 205 download: bool = False, 206 **kwargs 207) -> DataLoader: 208 """Get the RETOUCH dataloader for retinal fluid segmentation. 209 210 Args: 211 path: Filepath to a folder where the data is downloaded for further processing. 212 batch_size: The batch size for training. 213 patch_shape: The patch shape to use for training. 214 vendor: The choice of device vendor(s). By default, volumes from all vendors are used. 215 resize_inputs: Whether to resize inputs to the desired patch shape. 216 download: Whether to download the data if it is not present. 217 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 218 219 Returns: 220 The DataLoader. 221 """ 222 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 223 dataset = get_retouch_dataset(path, patch_shape, vendor, resize_inputs, download, **ds_kwargs) 224 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the RETOUCH dataloader for retinal fluid segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- vendor: The choice of device vendor(s). By default, volumes from all vendors are used.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.