torch_em.data.datasets.medical.rim_one_dl

The RIM-ONE DL dataset contains annotations for optic disc and optic cup segmentation in fundus images, for the task of glaucoma assessment.

The dataset is hosted at https://github.com/miag-ull/rim-one-dl. It comprises 485 retinographies (313 from normal subjects and 172 from patients with glaucoma), collected at three Spanish hospitals (Hospital Universitario de Canarias, Hospital Universitario Miguel Servet and Hospital Clinico Universitario San Carlos). The images and reference segmentations are distributed via the bit.ly links referenced in the repository's README (https://bit.ly/rim-one-dl-images and https://bit.ly/rim-one-dl-reference-segmentations), which resolve to Google Drive files; this module downloads directly from the resolved Google Drive files.

The dataset ships two official partitions, each with a training and a test set: 'random' (images distributed randomly between the two sets) and 'hospital' (the test set is built from two of the three hospitals, held out from the training set). The label masks are binary PNGs, one for the optic disc and one for the optic cup, produced with the DCSeg annotation tool by an expert in glaucoma.

The dataset is from the publication https://doi.org/10.5566/ias.2346. Please cite it if you use this dataset for your research. Data included in this database can only be used for research and educational purposes.

  1"""The RIM-ONE DL dataset contains annotations for optic disc and optic cup segmentation in fundus
  2images, for the task of glaucoma assessment.
  3
  4The dataset is hosted at https://github.com/miag-ull/rim-one-dl. It comprises 485 retinographies
  5(313 from normal subjects and 172 from patients with glaucoma), collected at three Spanish hospitals
  6(Hospital Universitario de Canarias, Hospital Universitario Miguel Servet and Hospital Clinico
  7Universitario San Carlos). The images and reference segmentations are distributed via the bit.ly
  8links referenced in the repository's README (https://bit.ly/rim-one-dl-images and
  9https://bit.ly/rim-one-dl-reference-segmentations), which resolve to Google Drive files; this module
 10downloads directly from the resolved Google Drive files.
 11
 12The dataset ships two official partitions, each with a training and a test set: 'random' (images
 13distributed randomly between the two sets) and 'hospital' (the test set is built from two of the
 14three hospitals, held out from the training set). The label masks are binary PNGs, one for the optic
 15disc and one for the optic cup, produced with the DCSeg annotation tool by an expert in glaucoma.
 16
 17The dataset is from the publication https://doi.org/10.5566/ias.2346.
 18Please cite it if you use this dataset for your research. Data included in this database can only be
 19used for research and educational purposes.
 20"""
 21
 22import os
 23from glob import glob
 24from pathlib import Path
 25from typing import Union, Tuple, Literal, List
 26
 27from torch.utils.data import Dataset, DataLoader
 28
 29import torch_em
 30
 31from .. import util
 32
 33
 34URL = {
 35    "images": "https://drive.google.com/uc?id=1teYi_smpLiNZNJcTWdxXgKKLW2fkUQr4",
 36    "segmentations": "https://drive.google.com/uc?id=1eb1V9V65TuwFNYmYsIzgdAgdWyD6o7bG",
 37}
 38
 39CHECKSUM = {
 40    "images": "85aed5f95c794f52d11b6ed953032108a66576ffcff4b587365df467491b603d",
 41    "segmentations": "edc363b1f0deabc8ed7a356250a9e1fb21825bf06f7337889fd1e764e1947901",
 42}
 43
 44PARTITION_DIRS = {"random": "partitioned_randomly", "hospital": "partitioned_by_hospital"}
 45SPLIT_DIRS = {"train": "training_set", "test": "test_set"}
 46TASK_NAMES = {"disc": "Disc", "cup": "Cup"}
 47
 48
 49def get_rim_one_dl_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 50    """Download the RIM-ONE DL dataset.
 51
 52    Args:
 53        path: Filepath to a folder where the data is downloaded for further processing.
 54        download: Whether to download the data if it is not present.
 55
 56    Returns:
 57        Filepath where the data is downloaded.
 58    """
 59    images_dir = os.path.join(path, "RIM-ONE_DL_images")
 60    if os.path.exists(images_dir):
 61        return path
 62
 63    os.makedirs(path, exist_ok=True)
 64
 65    zip_path = os.path.join(path, "rim_one_dl_images.zip")
 66    util.download_source_gdrive(
 67        path=zip_path, url=URL["images"], download=download, checksum=CHECKSUM["images"], download_type="zip",
 68    )
 69    util.unzip(zip_path=zip_path, dst=path)
 70
 71    zip_path = os.path.join(path, "rim_one_dl_segmentations.zip")
 72    util.download_source_gdrive(
 73        path=zip_path, url=URL["segmentations"], download=download, checksum=CHECKSUM["segmentations"],
 74        download_type="zip",
 75    )
 76    util.unzip(zip_path=zip_path, dst=path)
 77
 78    return path
 79
 80
 81def get_rim_one_dl_paths(
 82    path: Union[os.PathLike, str],
 83    split: Literal["train", "test"],
 84    partition: Literal["random", "hospital"] = "random",
 85    task: Literal["disc", "cup"] = "disc",
 86    download: bool = False,
 87) -> Tuple[List[str], List[str]]:
 88    """Get paths to the RIM-ONE DL data.
 89
 90    Args:
 91        path: Filepath to a folder where the data is downloaded for further processing.
 92        split: The choice of data split.
 93        partition: The choice of the official partition, either 'random' or 'hospital'.
 94        task: The choice of labels for the specific task.
 95        download: Whether to download the data if it is not present.
 96
 97    Returns:
 98        List of filepaths for the image data.
 99        List of filepaths for the label data.
100    """
101    root_dir = get_rim_one_dl_data(path=path, download=download)
102
103    assert split in SPLIT_DIRS, f"'{split}' is not a valid split."
104    assert partition in PARTITION_DIRS, f"'{partition}' is not a valid partition."
105    assert task in TASK_NAMES, f"'{task}' is not a valid task."
106
107    image_dir = os.path.join(root_dir, "RIM-ONE_DL_images", PARTITION_DIRS[partition], SPLIT_DIRS[split])
108    image_paths = sorted(glob(os.path.join(image_dir, "*", "*.png")))
109
110    gt_dir = os.path.join(root_dir, "RIM-ONE_DL_reference_segmentations")
111    gt_paths = [
112        os.path.join(gt_dir, Path(p).parent.name, f"{Path(p).stem}-1-{TASK_NAMES[task]}-T.png") for p in image_paths
113    ]
114
115    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
116    for gt_path in gt_paths:
117        assert os.path.exists(gt_path), gt_path
118
119    return image_paths, gt_paths
120
121
122def get_rim_one_dl_dataset(
123    path: Union[os.PathLike, str],
124    patch_shape: Tuple[int, int],
125    split: Literal["train", "test"],
126    partition: Literal["random", "hospital"] = "random",
127    task: Literal["disc", "cup"] = "disc",
128    resize_inputs: bool = False,
129    download: bool = False,
130    **kwargs
131) -> Dataset:
132    """Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images.
133
134    Args:
135        path: Filepath to a folder where the data is downloaded for further processing.
136        patch_shape: The patch shape to use for training.
137        split: The choice of data split.
138        partition: The choice of the official partition.
139        task: The choice of labels for the specific task.
140        resize_inputs: Whether to resize the inputs to the expected patch shape.
141        download: Whether to download the data if it is not present.
142        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
143
144    Returns:
145        The segmentation dataset.
146    """
147    image_paths, gt_paths = get_rim_one_dl_paths(path, split, partition, task, download)
148
149    if resize_inputs:
150        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
151        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
152            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
153        )
154
155    return torch_em.default_segmentation_dataset(
156        raw_paths=image_paths,
157        raw_key=None,
158        label_paths=gt_paths,
159        label_key=None,
160        patch_shape=patch_shape,
161        is_seg_dataset=False,
162        **kwargs
163    )
164
165
166def get_rim_one_dl_loader(
167    path: Union[os.PathLike, str],
168    batch_size: int,
169    patch_shape: Tuple[int, int],
170    split: Literal["train", "test"],
171    partition: Literal["random", "hospital"] = "random",
172    task: Literal["disc", "cup"] = "disc",
173    resize_inputs: bool = False,
174    download: bool = False,
175    **kwargs
176) -> DataLoader:
177    """Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images.
178
179    Args:
180        path: Filepath to a folder where the data is downloaded for further processing.
181        batch_size: The batch size for training.
182        patch_shape: The patch shape to use for training.
183        split: The choice of data split.
184        partition: The choice of the official partition.
185        task: The choice of labels for the specific task.
186        resize_inputs: Whether to resize the inputs to the expected patch shape.
187        download: Whether to download the data if it is not present.
188        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
189
190    Returns:
191        The DataLoader.
192    """
193    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
194    dataset = get_rim_one_dl_dataset(path, patch_shape, split, partition, task, resize_inputs, download, **ds_kwargs)
195    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = {'images': 'https://drive.google.com/uc?id=1teYi_smpLiNZNJcTWdxXgKKLW2fkUQr4', 'segmentations': 'https://drive.google.com/uc?id=1eb1V9V65TuwFNYmYsIzgdAgdWyD6o7bG'}
CHECKSUM = {'images': '85aed5f95c794f52d11b6ed953032108a66576ffcff4b587365df467491b603d', 'segmentations': 'edc363b1f0deabc8ed7a356250a9e1fb21825bf06f7337889fd1e764e1947901'}
PARTITION_DIRS = {'random': 'partitioned_randomly', 'hospital': 'partitioned_by_hospital'}
SPLIT_DIRS = {'train': 'training_set', 'test': 'test_set'}
TASK_NAMES = {'disc': 'Disc', 'cup': 'Cup'}
def get_rim_one_dl_data(path: Union[os.PathLike, str], download: bool = False) -> str:
50def get_rim_one_dl_data(path: Union[os.PathLike, str], download: bool = False) -> str:
51    """Download the RIM-ONE DL dataset.
52
53    Args:
54        path: Filepath to a folder where the data is downloaded for further processing.
55        download: Whether to download the data if it is not present.
56
57    Returns:
58        Filepath where the data is downloaded.
59    """
60    images_dir = os.path.join(path, "RIM-ONE_DL_images")
61    if os.path.exists(images_dir):
62        return path
63
64    os.makedirs(path, exist_ok=True)
65
66    zip_path = os.path.join(path, "rim_one_dl_images.zip")
67    util.download_source_gdrive(
68        path=zip_path, url=URL["images"], download=download, checksum=CHECKSUM["images"], download_type="zip",
69    )
70    util.unzip(zip_path=zip_path, dst=path)
71
72    zip_path = os.path.join(path, "rim_one_dl_segmentations.zip")
73    util.download_source_gdrive(
74        path=zip_path, url=URL["segmentations"], download=download, checksum=CHECKSUM["segmentations"],
75        download_type="zip",
76    )
77    util.unzip(zip_path=zip_path, dst=path)
78
79    return path

Download the RIM-ONE DL dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_rim_one_dl_paths( path: Union[os.PathLike, str], split: Literal['train', 'test'], partition: Literal['random', 'hospital'] = 'random', task: Literal['disc', 'cup'] = 'disc', download: bool = False) -> Tuple[List[str], List[str]]:
 82def get_rim_one_dl_paths(
 83    path: Union[os.PathLike, str],
 84    split: Literal["train", "test"],
 85    partition: Literal["random", "hospital"] = "random",
 86    task: Literal["disc", "cup"] = "disc",
 87    download: bool = False,
 88) -> Tuple[List[str], List[str]]:
 89    """Get paths to the RIM-ONE DL data.
 90
 91    Args:
 92        path: Filepath to a folder where the data is downloaded for further processing.
 93        split: The choice of data split.
 94        partition: The choice of the official partition, either 'random' or 'hospital'.
 95        task: The choice of labels for the specific task.
 96        download: Whether to download the data if it is not present.
 97
 98    Returns:
 99        List of filepaths for the image data.
100        List of filepaths for the label data.
101    """
102    root_dir = get_rim_one_dl_data(path=path, download=download)
103
104    assert split in SPLIT_DIRS, f"'{split}' is not a valid split."
105    assert partition in PARTITION_DIRS, f"'{partition}' is not a valid partition."
106    assert task in TASK_NAMES, f"'{task}' is not a valid task."
107
108    image_dir = os.path.join(root_dir, "RIM-ONE_DL_images", PARTITION_DIRS[partition], SPLIT_DIRS[split])
109    image_paths = sorted(glob(os.path.join(image_dir, "*", "*.png")))
110
111    gt_dir = os.path.join(root_dir, "RIM-ONE_DL_reference_segmentations")
112    gt_paths = [
113        os.path.join(gt_dir, Path(p).parent.name, f"{Path(p).stem}-1-{TASK_NAMES[task]}-T.png") for p in image_paths
114    ]
115
116    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
117    for gt_path in gt_paths:
118        assert os.path.exists(gt_path), gt_path
119
120    return image_paths, gt_paths

Get paths to the RIM-ONE DL data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • split: The choice of data split.
  • partition: The choice of the official partition, either 'random' or 'hospital'.
  • task: The choice of labels for the specific task.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_rim_one_dl_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], split: Literal['train', 'test'], partition: Literal['random', 'hospital'] = 'random', task: Literal['disc', 'cup'] = 'disc', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
123def get_rim_one_dl_dataset(
124    path: Union[os.PathLike, str],
125    patch_shape: Tuple[int, int],
126    split: Literal["train", "test"],
127    partition: Literal["random", "hospital"] = "random",
128    task: Literal["disc", "cup"] = "disc",
129    resize_inputs: bool = False,
130    download: bool = False,
131    **kwargs
132) -> Dataset:
133    """Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images.
134
135    Args:
136        path: Filepath to a folder where the data is downloaded for further processing.
137        patch_shape: The patch shape to use for training.
138        split: The choice of data split.
139        partition: The choice of the official partition.
140        task: The choice of labels for the specific task.
141        resize_inputs: Whether to resize the inputs to the expected patch shape.
142        download: Whether to download the data if it is not present.
143        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
144
145    Returns:
146        The segmentation dataset.
147    """
148    image_paths, gt_paths = get_rim_one_dl_paths(path, split, partition, task, download)
149
150    if resize_inputs:
151        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
152        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
153            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
154        )
155
156    return torch_em.default_segmentation_dataset(
157        raw_paths=image_paths,
158        raw_key=None,
159        label_paths=gt_paths,
160        label_key=None,
161        patch_shape=patch_shape,
162        is_seg_dataset=False,
163        **kwargs
164    )

Get the RIM-ONE DL dataset for segmentation of optic disc and optic cup in fundus images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • split: The choice of data split.
  • partition: The choice of the official partition.
  • task: The choice of labels for the specific task.
  • resize_inputs: Whether to resize the inputs to the expected patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_rim_one_dl_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], split: Literal['train', 'test'], partition: Literal['random', 'hospital'] = 'random', task: Literal['disc', 'cup'] = 'disc', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
167def get_rim_one_dl_loader(
168    path: Union[os.PathLike, str],
169    batch_size: int,
170    patch_shape: Tuple[int, int],
171    split: Literal["train", "test"],
172    partition: Literal["random", "hospital"] = "random",
173    task: Literal["disc", "cup"] = "disc",
174    resize_inputs: bool = False,
175    download: bool = False,
176    **kwargs
177) -> DataLoader:
178    """Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images.
179
180    Args:
181        path: Filepath to a folder where the data is downloaded for further processing.
182        batch_size: The batch size for training.
183        patch_shape: The patch shape to use for training.
184        split: The choice of data split.
185        partition: The choice of the official partition.
186        task: The choice of labels for the specific task.
187        resize_inputs: Whether to resize the inputs to the expected patch shape.
188        download: Whether to download the data if it is not present.
189        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
190
191    Returns:
192        The DataLoader.
193    """
194    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
195    dataset = get_rim_one_dl_dataset(path, patch_shape, split, partition, task, resize_inputs, download, **ds_kwargs)
196    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the RIM-ONE DL dataloader for segmentation of optic disc and optic cup in fundus images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • split: The choice of data split.
  • partition: The choice of the official partition.
  • task: The choice of labels for the specific task.
  • resize_inputs: Whether to resize the inputs to the expected patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.