torch_em.data.datasets.medical.rim_one_dl
The RIM-ONE DL dataset contains annotations for optic disc and optic cup segmentation in fundus images, for the task of glaucoma assessment.
The dataset is hosted at https://github.com/miag-ull/rim-one-dl. It comprises 485 retinographies (313 from normal subjects and 172 from patients with glaucoma), collected at three Spanish hospitals (Hospital Universitario de Canarias, Hospital Universitario Miguel Servet and Hospital Clinico Universitario San Carlos). The images and reference segmentations are distributed via the bit.ly links referenced in the repository's README (https://bit.ly/rim-one-dl-images and https://bit.ly/rim-one-dl-reference-segmentations), which resolve to Google Drive files; this module downloads directly from the resolved Google Drive files.
The dataset ships two official partitions, each with a training and a test set: 'random' (images distributed randomly between the two sets) and 'hospital' (the test set is built from two of the three hospitals, held out from the training set). The label masks are binary PNGs, one for the optic disc and one for the optic cup, produced with the DCSeg annotation tool by an expert in glaucoma.
The dataset is from the publication https://doi.org/10.5566/ias.2346. Please cite it if you use this dataset for your research. Data included in this database can only be used for research and educational purposes.
1"""The RIM-ONE DL dataset contains annotations for optic disc and optic cup segmentation in fundus 2images, for the task of glaucoma assessment. 3 4The dataset is hosted at https://github.com/miag-ull/rim-one-dl. It comprises 485 retinographies 5(313 from normal subjects and 172 from patients with glaucoma), collected at three Spanish hospitals 6(Hospital Universitario de Canarias, Hospital Universitario Miguel Servet and Hospital Clinico 7Universitario San Carlos). The images and reference segmentations are distributed via the bit.ly 8links referenced in the repository's README (https://bit.ly/rim-one-dl-images and 9https://bit.ly/rim-one-dl-reference-segmentations), which resolve to Google Drive files; this module 10downloads directly from the resolved Google Drive files. 11 12The dataset ships two official partitions, each with a training and a test set: 'random' (images 13distributed randomly between the two sets) and 'hospital' (the test set is built from two of the 14three hospitals, held out from the training set). The label masks are binary PNGs, one for the optic 15disc and one for the optic cup, produced with the DCSeg annotation tool by an expert in glaucoma. 16 17The dataset is from the publication https://doi.org/10.5566/ias.2346. 18Please cite it if you use this dataset for your research. Data included in this database can only be 19used for research and educational purposes. 20""" 21 22import os 23from glob import glob 24from pathlib import Path 25from typing import Union, Tuple, Literal, List 26 27from torch.utils.data import Dataset, DataLoader 28 29import torch_em 30 31from .. import util 32 33 34URL = { 35 "images": "https://drive.google.com/uc?id=1teYi_smpLiNZNJcTWdxXgKKLW2fkUQr4", 36 "segmentations": "https://drive.google.com/uc?id=1eb1V9V65TuwFNYmYsIzgdAgdWyD6o7bG", 37} 38 39CHECKSUM = { 40 "images": "85aed5f95c794f52d11b6ed953032108a66576ffcff4b587365df467491b603d", 41 "segmentations": "edc363b1f0deabc8ed7a356250a9e1fb21825bf06f7337889fd1e764e1947901", 42} 43 44PARTITION_DIRS = {"random": "partitioned_randomly", "hospital": "partitioned_by_hospital"} 45SPLIT_DIRS = {"train": "training_set", "test": "test_set"} 46TASK_NAMES = {"disc": "Disc", "cup": "Cup"} 47 48 49def get_rim_one_dl_data(path: Union[os.PathLike, str], download: bool = False) -> str: 50 """Download the RIM-ONE DL dataset. 51 52 Args: 53 path: Filepath to a folder where the data is downloaded for further processing. 54 download: Whether to download the data if it is not present. 55 56 Returns: 57 Filepath where the data is downloaded. 58 """ 59 images_dir = os.path.join(path, "RIM-ONE_DL_images") 60 if os.path.exists(images_dir): 61 return path 62 63 os.makedirs(path, exist_ok=True) 64 65 zip_path = os.path.join(path, "rim_one_dl_images.zip") 66 util.download_source_gdrive( 67 path=zip_path, url=URL["images"], download=download, checksum=CHECKSUM["images"], download_type="zip", 68 ) 69 util.unzip(zip_path=zip_path, dst=path) 70 71 zip_path = os.path.join(path, "rim_one_dl_segmentations.zip") 72 util.download_source_gdrive( 73 path=zip_path, url=URL["segmentations"], download=download, checksum=CHECKSUM["segmentations"], 74 download_type="zip", 75 ) 76 util.unzip(zip_path=zip_path, dst=path) 77 78 return path 79 80 81def get_rim_one_dl_paths( 82 path: Union[os.PathLike, str], 83 split: Literal["train", "test"], 84 partition: Literal["random", "hospital"] = "random", 85 task: Literal["disc", "cup"] = "disc", 86 download: bool = False, 87) -> Tuple[List[str], List[str]]: 88 """Get paths to the RIM-ONE DL data. 89 90 Args: 91 path: Filepath to a folder where the data is downloaded for further processing. 92 split: The choice of data split. 93 partition: The choice of the official partition, either 'random' or 'hospital'. 94 task: The choice of labels for the specific task. 95 download: Whether to download the data if it is not present. 96 97 Returns: 98 List of filepaths for the image data. 99 List of filepaths for the label data. 100 """ 101 root_dir = get_rim_one_dl_data(path=path, download=download) 102 103 assert split in SPLIT_DIRS, f"'{split}' is not a valid split." 104 assert partition in PARTITION_DIRS, f"'{partition}' is not a valid partition." 105 assert task in TASK_NAMES, f"'{task}' is not a valid task." 106 107 image_dir = os.path.join(root_dir, "RIM-ONE_DL_images", PARTITION_DIRS[partition], SPLIT_DIRS[split]) 108 image_paths = sorted(glob(os.path.join(image_dir, "*", "*.png"))) 109 110 gt_dir = os.path.join(root_dir, "RIM-ONE_DL_reference_segmentations") 111 gt_paths = [ 112 os.path.join(gt_dir, Path(p).parent.name, f"{Path(p).stem}-1-{TASK_NAMES[task]}-T.png") for p in image_paths 113 ] 114 115 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 116 for gt_path in gt_paths: 117 assert os.path.exists(gt_path), gt_path 118 119 return image_paths, gt_paths 120 121 122def get_rim_one_dl_dataset( 123 path: Union[os.PathLike, str], 124 patch_shape: Tuple[int, int], 125 split: Literal["train", "test"], 126 partition: Literal["random", "hospital"] = "random", 127 task: Literal["disc", "cup"] = "disc", 128 resize_inputs: bool = False, 129 download: bool = False, 130 **kwargs 131) -> Dataset: 132 """Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images. 133 134 Args: 135 path: Filepath to a folder where the data is downloaded for further processing. 136 patch_shape: The patch shape to use for training. 137 split: The choice of data split. 138 partition: The choice of the official partition. 139 task: The choice of labels for the specific task. 140 resize_inputs: Whether to resize the inputs to the expected patch shape. 141 download: Whether to download the data if it is not present. 142 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 143 144 Returns: 145 The segmentation dataset. 146 """ 147 image_paths, gt_paths = get_rim_one_dl_paths(path, split, partition, task, download) 148 149 if resize_inputs: 150 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 151 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 152 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 153 ) 154 155 return torch_em.default_segmentation_dataset( 156 raw_paths=image_paths, 157 raw_key=None, 158 label_paths=gt_paths, 159 label_key=None, 160 patch_shape=patch_shape, 161 is_seg_dataset=False, 162 **kwargs 163 ) 164 165 166def get_rim_one_dl_loader( 167 path: Union[os.PathLike, str], 168 batch_size: int, 169 patch_shape: Tuple[int, int], 170 split: Literal["train", "test"], 171 partition: Literal["random", "hospital"] = "random", 172 task: Literal["disc", "cup"] = "disc", 173 resize_inputs: bool = False, 174 download: bool = False, 175 **kwargs 176) -> DataLoader: 177 """Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images. 178 179 Args: 180 path: Filepath to a folder where the data is downloaded for further processing. 181 batch_size: The batch size for training. 182 patch_shape: The patch shape to use for training. 183 split: The choice of data split. 184 partition: The choice of the official partition. 185 task: The choice of labels for the specific task. 186 resize_inputs: Whether to resize the inputs to the expected patch shape. 187 download: Whether to download the data if it is not present. 188 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 189 190 Returns: 191 The DataLoader. 192 """ 193 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 194 dataset = get_rim_one_dl_dataset(path, patch_shape, split, partition, task, resize_inputs, download, **ds_kwargs) 195 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
50def get_rim_one_dl_data(path: Union[os.PathLike, str], download: bool = False) -> str: 51 """Download the RIM-ONE DL dataset. 52 53 Args: 54 path: Filepath to a folder where the data is downloaded for further processing. 55 download: Whether to download the data if it is not present. 56 57 Returns: 58 Filepath where the data is downloaded. 59 """ 60 images_dir = os.path.join(path, "RIM-ONE_DL_images") 61 if os.path.exists(images_dir): 62 return path 63 64 os.makedirs(path, exist_ok=True) 65 66 zip_path = os.path.join(path, "rim_one_dl_images.zip") 67 util.download_source_gdrive( 68 path=zip_path, url=URL["images"], download=download, checksum=CHECKSUM["images"], download_type="zip", 69 ) 70 util.unzip(zip_path=zip_path, dst=path) 71 72 zip_path = os.path.join(path, "rim_one_dl_segmentations.zip") 73 util.download_source_gdrive( 74 path=zip_path, url=URL["segmentations"], download=download, checksum=CHECKSUM["segmentations"], 75 download_type="zip", 76 ) 77 util.unzip(zip_path=zip_path, dst=path) 78 79 return path
Download the RIM-ONE DL dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is downloaded.
82def get_rim_one_dl_paths( 83 path: Union[os.PathLike, str], 84 split: Literal["train", "test"], 85 partition: Literal["random", "hospital"] = "random", 86 task: Literal["disc", "cup"] = "disc", 87 download: bool = False, 88) -> Tuple[List[str], List[str]]: 89 """Get paths to the RIM-ONE DL data. 90 91 Args: 92 path: Filepath to a folder where the data is downloaded for further processing. 93 split: The choice of data split. 94 partition: The choice of the official partition, either 'random' or 'hospital'. 95 task: The choice of labels for the specific task. 96 download: Whether to download the data if it is not present. 97 98 Returns: 99 List of filepaths for the image data. 100 List of filepaths for the label data. 101 """ 102 root_dir = get_rim_one_dl_data(path=path, download=download) 103 104 assert split in SPLIT_DIRS, f"'{split}' is not a valid split." 105 assert partition in PARTITION_DIRS, f"'{partition}' is not a valid partition." 106 assert task in TASK_NAMES, f"'{task}' is not a valid task." 107 108 image_dir = os.path.join(root_dir, "RIM-ONE_DL_images", PARTITION_DIRS[partition], SPLIT_DIRS[split]) 109 image_paths = sorted(glob(os.path.join(image_dir, "*", "*.png"))) 110 111 gt_dir = os.path.join(root_dir, "RIM-ONE_DL_reference_segmentations") 112 gt_paths = [ 113 os.path.join(gt_dir, Path(p).parent.name, f"{Path(p).stem}-1-{TASK_NAMES[task]}-T.png") for p in image_paths 114 ] 115 116 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 117 for gt_path in gt_paths: 118 assert os.path.exists(gt_path), gt_path 119 120 return image_paths, gt_paths
Get paths to the RIM-ONE DL data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- split: The choice of data split.
- partition: The choice of the official partition, either 'random' or 'hospital'.
- task: The choice of labels for the specific task.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
123def get_rim_one_dl_dataset( 124 path: Union[os.PathLike, str], 125 patch_shape: Tuple[int, int], 126 split: Literal["train", "test"], 127 partition: Literal["random", "hospital"] = "random", 128 task: Literal["disc", "cup"] = "disc", 129 resize_inputs: bool = False, 130 download: bool = False, 131 **kwargs 132) -> Dataset: 133 """Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images. 134 135 Args: 136 path: Filepath to a folder where the data is downloaded for further processing. 137 patch_shape: The patch shape to use for training. 138 split: The choice of data split. 139 partition: The choice of the official partition. 140 task: The choice of labels for the specific task. 141 resize_inputs: Whether to resize the inputs to the expected patch shape. 142 download: Whether to download the data if it is not present. 143 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 144 145 Returns: 146 The segmentation dataset. 147 """ 148 image_paths, gt_paths = get_rim_one_dl_paths(path, split, partition, task, download) 149 150 if resize_inputs: 151 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 152 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 153 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 154 ) 155 156 return torch_em.default_segmentation_dataset( 157 raw_paths=image_paths, 158 raw_key=None, 159 label_paths=gt_paths, 160 label_key=None, 161 patch_shape=patch_shape, 162 is_seg_dataset=False, 163 **kwargs 164 )
Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- split: The choice of data split.
- partition: The choice of the official partition.
- task: The choice of labels for the specific task.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
167def get_rim_one_dl_loader( 168 path: Union[os.PathLike, str], 169 batch_size: int, 170 patch_shape: Tuple[int, int], 171 split: Literal["train", "test"], 172 partition: Literal["random", "hospital"] = "random", 173 task: Literal["disc", "cup"] = "disc", 174 resize_inputs: bool = False, 175 download: bool = False, 176 **kwargs 177) -> DataLoader: 178 """Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images. 179 180 Args: 181 path: Filepath to a folder where the data is downloaded for further processing. 182 batch_size: The batch size for training. 183 patch_shape: The patch shape to use for training. 184 split: The choice of data split. 185 partition: The choice of the official partition. 186 task: The choice of labels for the specific task. 187 resize_inputs: Whether to resize the inputs to the expected patch shape. 188 download: Whether to download the data if it is not present. 189 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 190 191 Returns: 192 The DataLoader. 193 """ 194 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 195 dataset = get_rim_one_dl_dataset(path, patch_shape, split, partition, task, resize_inputs, download, **ds_kwargs) 196 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- split: The choice of data split.
- partition: The choice of the official partition.
- task: The choice of labels for the specific task.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.