torch_em.data.datasets.medical.octa500
The OCTA-500 dataset contains annotations for retinal vessel segmentation (large vessels, capillaries, arteries, veins) and foveal avascular zone (FAZ) segmentation in en-face projections derived from 3D OCT / OCTA volumes.
The dataset comprises 500 subjects, split into two field-of-view (FOV) subsets: 'OCTA_6M' (subject ids 10001-10300, FOV 6mm x 6mm x 2mm, volume shape 400 x 400 x 640) and 'OCTA_3M' (subject ids 10301-10500, FOV 3mm x 3mm x 2mm, volume shape 304 x 304 x 640). For each subject, six 2D en-face projection maps are derived from the 3D OCT / OCTA volumes: 'OCT(FULL)', 'OCT(ILM_OPL)', 'OCT(OPL_BM)', 'OCTA(FULL)', 'OCTA(ILM_OPL)' and 'OCTA(OPL_BM)'. Pixel-wise segmentation masks are provided for large vessels, capillaries, arteries, veins and the FAZ, matching the resolution of the projection maps.
NOTE: The dataset also provides 2D / 3D FAZ and retinal layer annotations in the original release, but
this loader only covers the vessel and FAZ label types listed in LABEL_DIRS above, whose folder layout
and file format could be corroborated from the dataset's accompanying publications and downstream usage.
The retinal layer annotations are not supported here, as their exact on-disk format could not be verified
without direct access to the (gated) data.
NOTE: This dataset is hosted on IEEE DataPort at https://ieee-dataport.org/open-access/octa-500 and is
gated: downloading it requires a free IEEE account (or IEEE Society membership) to view the page, and
the password-protected archives themselves require directly emailing the dataset's authors. Automatic
download is not supported, see get_octa500_data for the exact manual steps.
The dataset is from the publication https://doi.org/10.1016/j.media.2024.103092 (and the earlier preprint https://doi.org/10.48550/arXiv.2012.07261). Please cite it if you use this dataset in your research.
1"""The OCTA-500 dataset contains annotations for retinal vessel segmentation (large vessels, capillaries, 2arteries, veins) and foveal avascular zone (FAZ) segmentation in en-face projections derived from 3D 3OCT / OCTA volumes. 4 5The dataset comprises 500 subjects, split into two field-of-view (FOV) subsets: 'OCTA_6M' (subject ids 610001-10300, FOV 6mm x 6mm x 2mm, volume shape 400 x 400 x 640) and 'OCTA_3M' (subject ids 10301-10500, 7FOV 3mm x 3mm x 2mm, volume shape 304 x 304 x 640). For each subject, six 2D en-face projection maps are 8derived from the 3D OCT / OCTA volumes: 'OCT(FULL)', 'OCT(ILM_OPL)', 'OCT(OPL_BM)', 'OCTA(FULL)', 9'OCTA(ILM_OPL)' and 'OCTA(OPL_BM)'. Pixel-wise segmentation masks are provided for large vessels, 10capillaries, arteries, veins and the FAZ, matching the resolution of the projection maps. 11 12NOTE: The dataset also provides 2D / 3D FAZ and retinal layer annotations in the original release, but 13this loader only covers the vessel and FAZ label types listed in `LABEL_DIRS` above, whose folder layout 14and file format could be corroborated from the dataset's accompanying publications and downstream usage. 15The retinal layer annotations are not supported here, as their exact on-disk format could not be verified 16without direct access to the (gated) data. 17 18NOTE: This dataset is hosted on IEEE DataPort at https://ieee-dataport.org/open-access/octa-500 and is 19gated: downloading it requires a free IEEE account (or IEEE Society membership) to view the page, and 20the password-protected archives themselves require directly emailing the dataset's authors. Automatic 21download is not supported, see `get_octa500_data` for the exact manual steps. 22 23The dataset is from the publication https://doi.org/10.1016/j.media.2024.103092 (and the earlier preprint 24https://doi.org/10.48550/arXiv.2012.07261). Please cite it if you use this dataset in your research. 25""" 26 27import os 28from glob import glob 29from natsort import natsorted 30from typing import Union, Tuple, Literal, List 31 32from torch.utils.data import Dataset, DataLoader 33 34import torch_em 35 36from .. import util 37 38 39SUBSET_IDS = { 40 "6M": range(10001, 10301), 41 "3M": range(10301, 10501), 42} 43"""Mapping from the subset choice to its range of subject ids.""" 44 45PROJECTIONS = { 46 "oct_full": "OCT(FULL)", 47 "oct_ilm_opl": "OCT(ILM_OPL)", 48 "oct_opl_bm": "OCT(OPL_BM)", 49 "octa_full": "OCTA(FULL)", 50 "octa_ilm_opl": "OCTA(ILM_OPL)", 51 "octa_opl_bm": "OCTA(OPL_BM)", 52} 53"""Mapping from the projection choice to its folder in the released data.""" 54 55LABEL_DIRS = { 56 "large_vessel": "GT_LargeVessel", 57 "capillary": "GT_Capillary", 58 "artery": "GT_Artery", 59 "vein": "GT_Vein", 60 "faz": "GT_FAZ", 61} 62"""Mapping from the label type choice to its folder in the released data.""" 63 64 65def get_octa500_data( 66 path: Union[os.PathLike, str], subset: Literal["3M", "6M"], download: bool = False 67) -> str: 68 """Obtain the OCTA-500 dataset. 69 70 Args: 71 path: Filepath to a folder where the data is downloaded for further processing. 72 subset: The choice of field-of-view subset. Either '3M' or '6M'. 73 download: Whether to download the data if it is not present. 74 75 Returns: 76 Filepath to the folder where the subset data is expected to be stored. 77 """ 78 if subset not in SUBSET_IDS: 79 raise ValueError(f"'{subset}' is not a valid subset. Choose from {list(SUBSET_IDS.keys())}.") 80 81 data_dir = os.path.join(path, f"OCTA_{subset}") 82 if os.path.exists(data_dir): 83 return data_dir 84 85 if download: 86 msg = "Download is set to True, but 'torch_em' cannot download this dataset automatically." 87 raise NotImplementedError(msg) 88 89 raise RuntimeError( 90 "The OCTA-500 dataset is hosted on IEEE DataPort and cannot be downloaded automatically. " 91 "Please follow these steps to obtain it manually:\n" 92 "1. Visit https://ieee-dataport.org/open-access/octa-500 and sign in with a free IEEE account " 93 "(or IEEE Society membership), which is required to view the download page.\n" 94 "2. The archives are password-protected. Send an email to chen2qiang@njust.edu.cn with the " 95 "subject line 'OCTA500: [your organization]: [your name]' to request the password.\n" 96 "3. Once you have the password, download 'Label.zip' and the 'OCTA_3mm_part*.zip' / " 97 "'OCTA_6mm_part*.zip' archives from the IEEE DataPort page and extract them.\n" 98 f"4. Place (or symlink) the extracted '3M' subset at '{os.path.join(path, 'OCTA_3M')}' and the " 99 f"'6M' subset at '{os.path.join(path, 'OCTA_6M')}', each containing a 'Projection Maps' folder " 100 "and the 'GT_LargeVessel' / 'GT_Capillary' / 'GT_Artery' / 'GT_Vein' / 'GT_FAZ' label folders.\n" 101 f"Expected location for the requested '{subset}' subset: '{data_dir}'." 102 ) 103 104 105def get_octa500_paths( 106 path: Union[os.PathLike, str], 107 subset: Literal["3M", "6M"], 108 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 109 projection: Literal[ 110 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 111 ] = "octa_full", 112 download: bool = False, 113) -> Tuple[List[str], List[str]]: 114 """Get paths to the OCTA-500 data. 115 116 Args: 117 path: Filepath to a folder where the data is downloaded for further processing. 118 subset: The choice of field-of-view subset. Either '3M' or '6M'. 119 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 120 'vein' or 'faz'. 121 projection: The choice of 2D en-face projection map used as the raw input. 122 download: Whether to download the data if it is not present. 123 124 Returns: 125 List of filepaths for the image data. 126 List of filepaths for the label data. 127 """ 128 if label_type not in LABEL_DIRS: 129 raise ValueError(f"'{label_type}' is not a valid label type. Choose from {list(LABEL_DIRS.keys())}.") 130 131 if projection not in PROJECTIONS: 132 raise ValueError(f"'{projection}' is not a valid projection. Choose from {list(PROJECTIONS.keys())}.") 133 134 data_dir = get_octa500_data(path, subset, download) 135 136 image_dir = os.path.join(data_dir, "Projection Maps", PROJECTIONS[projection]) 137 label_dir = os.path.join(data_dir, LABEL_DIRS[label_type]) 138 139 image_paths, label_paths = [], [] 140 for subject_id in SUBSET_IDS[subset]: 141 label_matches = glob(os.path.join(label_dir, f"{subject_id}.*")) 142 if not label_matches: 143 continue 144 145 image_matches = glob(os.path.join(image_dir, f"{subject_id}.*")) 146 if not image_matches: 147 continue 148 149 image_paths.append(image_matches[0]) 150 label_paths.append(label_matches[0]) 151 152 image_paths, label_paths = natsorted(image_paths), natsorted(label_paths) 153 assert len(image_paths) == len(label_paths) and len(image_paths) > 0, \ 154 f"Could not find matching image and '{label_type}' label pairs for the '{subset}' subset in '{data_dir}'." 155 156 return image_paths, label_paths 157 158 159def get_octa500_dataset( 160 path: Union[os.PathLike, str], 161 patch_shape: Tuple[int, int], 162 subset: Literal["3M", "6M"], 163 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 164 projection: Literal[ 165 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 166 ] = "octa_full", 167 resize_inputs: bool = False, 168 download: bool = False, 169 **kwargs 170) -> Dataset: 171 """Get the OCTA-500 dataset for retinal vessel and FAZ segmentation in OCTA en-face projections. 172 173 Args: 174 path: Filepath to a folder where the data is downloaded for further processing. 175 patch_shape: The patch shape to use for training. 176 subset: The choice of field-of-view subset. Either '3M' or '6M'. 177 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 178 'vein' or 'faz'. 179 projection: The choice of 2D en-face projection map used as the raw input. 180 resize_inputs: Whether to resize the inputs to the expected patch shape. 181 download: Whether to download the data if it is not present. 182 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 183 184 Returns: 185 The segmentation dataset. 186 """ 187 image_paths, label_paths = get_octa500_paths(path, subset, label_type, projection, download) 188 189 if resize_inputs: 190 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 191 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 192 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 193 ) 194 195 return torch_em.default_segmentation_dataset( 196 raw_paths=image_paths, 197 raw_key=None, 198 label_paths=label_paths, 199 label_key=None, 200 patch_shape=patch_shape, 201 is_seg_dataset=False, 202 **kwargs 203 ) 204 205 206def get_octa500_loader( 207 path: Union[os.PathLike, str], 208 batch_size: int, 209 patch_shape: Tuple[int, int], 210 subset: Literal["3M", "6M"], 211 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 212 projection: Literal[ 213 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 214 ] = "octa_full", 215 resize_inputs: bool = False, 216 download: bool = False, 217 **kwargs 218) -> DataLoader: 219 """Get the OCTA-500 dataloader for retinal vessel and FAZ segmentation in OCTA en-face projections. 220 221 Args: 222 path: Filepath to a folder where the data is downloaded for further processing. 223 batch_size: The batch size for training. 224 patch_shape: The patch shape to use for training. 225 subset: The choice of field-of-view subset. Either '3M' or '6M'. 226 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 227 'vein' or 'faz'. 228 projection: The choice of 2D en-face projection map used as the raw input. 229 resize_inputs: Whether to resize the inputs to the expected patch shape. 230 download: Whether to download the data if it is not present. 231 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 232 233 Returns: 234 The DataLoader. 235 """ 236 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 237 dataset = get_octa500_dataset( 238 path, patch_shape, subset, label_type, projection, resize_inputs, download, **ds_kwargs 239 ) 240 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Mapping from the subset choice to its range of subject ids.
Mapping from the projection choice to its folder in the released data.
Mapping from the label type choice to its folder in the released data.
66def get_octa500_data( 67 path: Union[os.PathLike, str], subset: Literal["3M", "6M"], download: bool = False 68) -> str: 69 """Obtain the OCTA-500 dataset. 70 71 Args: 72 path: Filepath to a folder where the data is downloaded for further processing. 73 subset: The choice of field-of-view subset. Either '3M' or '6M'. 74 download: Whether to download the data if it is not present. 75 76 Returns: 77 Filepath to the folder where the subset data is expected to be stored. 78 """ 79 if subset not in SUBSET_IDS: 80 raise ValueError(f"'{subset}' is not a valid subset. Choose from {list(SUBSET_IDS.keys())}.") 81 82 data_dir = os.path.join(path, f"OCTA_{subset}") 83 if os.path.exists(data_dir): 84 return data_dir 85 86 if download: 87 msg = "Download is set to True, but 'torch_em' cannot download this dataset automatically." 88 raise NotImplementedError(msg) 89 90 raise RuntimeError( 91 "The OCTA-500 dataset is hosted on IEEE DataPort and cannot be downloaded automatically. " 92 "Please follow these steps to obtain it manually:\n" 93 "1. Visit https://ieee-dataport.org/open-access/octa-500 and sign in with a free IEEE account " 94 "(or IEEE Society membership), which is required to view the download page.\n" 95 "2. The archives are password-protected. Send an email to chen2qiang@njust.edu.cn with the " 96 "subject line 'OCTA500: [your organization]: [your name]' to request the password.\n" 97 "3. Once you have the password, download 'Label.zip' and the 'OCTA_3mm_part*.zip' / " 98 "'OCTA_6mm_part*.zip' archives from the IEEE DataPort page and extract them.\n" 99 f"4. Place (or symlink) the extracted '3M' subset at '{os.path.join(path, 'OCTA_3M')}' and the " 100 f"'6M' subset at '{os.path.join(path, 'OCTA_6M')}', each containing a 'Projection Maps' folder " 101 "and the 'GT_LargeVessel' / 'GT_Capillary' / 'GT_Artery' / 'GT_Vein' / 'GT_FAZ' label folders.\n" 102 f"Expected location for the requested '{subset}' subset: '{data_dir}'." 103 )
Obtain the OCTA-500 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- subset: The choice of field-of-view subset. Either '3M' or '6M'.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the subset data is expected to be stored.
106def get_octa500_paths( 107 path: Union[os.PathLike, str], 108 subset: Literal["3M", "6M"], 109 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 110 projection: Literal[ 111 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 112 ] = "octa_full", 113 download: bool = False, 114) -> Tuple[List[str], List[str]]: 115 """Get paths to the OCTA-500 data. 116 117 Args: 118 path: Filepath to a folder where the data is downloaded for further processing. 119 subset: The choice of field-of-view subset. Either '3M' or '6M'. 120 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 121 'vein' or 'faz'. 122 projection: The choice of 2D en-face projection map used as the raw input. 123 download: Whether to download the data if it is not present. 124 125 Returns: 126 List of filepaths for the image data. 127 List of filepaths for the label data. 128 """ 129 if label_type not in LABEL_DIRS: 130 raise ValueError(f"'{label_type}' is not a valid label type. Choose from {list(LABEL_DIRS.keys())}.") 131 132 if projection not in PROJECTIONS: 133 raise ValueError(f"'{projection}' is not a valid projection. Choose from {list(PROJECTIONS.keys())}.") 134 135 data_dir = get_octa500_data(path, subset, download) 136 137 image_dir = os.path.join(data_dir, "Projection Maps", PROJECTIONS[projection]) 138 label_dir = os.path.join(data_dir, LABEL_DIRS[label_type]) 139 140 image_paths, label_paths = [], [] 141 for subject_id in SUBSET_IDS[subset]: 142 label_matches = glob(os.path.join(label_dir, f"{subject_id}.*")) 143 if not label_matches: 144 continue 145 146 image_matches = glob(os.path.join(image_dir, f"{subject_id}.*")) 147 if not image_matches: 148 continue 149 150 image_paths.append(image_matches[0]) 151 label_paths.append(label_matches[0]) 152 153 image_paths, label_paths = natsorted(image_paths), natsorted(label_paths) 154 assert len(image_paths) == len(label_paths) and len(image_paths) > 0, \ 155 f"Could not find matching image and '{label_type}' label pairs for the '{subset}' subset in '{data_dir}'." 156 157 return image_paths, label_paths
Get paths to the OCTA-500 data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- subset: The choice of field-of-view subset. Either '3M' or '6M'.
- label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 'vein' or 'faz'.
- projection: The choice of 2D en-face projection map used as the raw input.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
160def get_octa500_dataset( 161 path: Union[os.PathLike, str], 162 patch_shape: Tuple[int, int], 163 subset: Literal["3M", "6M"], 164 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 165 projection: Literal[ 166 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 167 ] = "octa_full", 168 resize_inputs: bool = False, 169 download: bool = False, 170 **kwargs 171) -> Dataset: 172 """Get the OCTA-500 dataset for retinal vessel and FAZ segmentation in OCTA en-face projections. 173 174 Args: 175 path: Filepath to a folder where the data is downloaded for further processing. 176 patch_shape: The patch shape to use for training. 177 subset: The choice of field-of-view subset. Either '3M' or '6M'. 178 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 179 'vein' or 'faz'. 180 projection: The choice of 2D en-face projection map used as the raw input. 181 resize_inputs: Whether to resize the inputs to the expected patch shape. 182 download: Whether to download the data if it is not present. 183 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 184 185 Returns: 186 The segmentation dataset. 187 """ 188 image_paths, label_paths = get_octa500_paths(path, subset, label_type, projection, download) 189 190 if resize_inputs: 191 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 192 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 193 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 194 ) 195 196 return torch_em.default_segmentation_dataset( 197 raw_paths=image_paths, 198 raw_key=None, 199 label_paths=label_paths, 200 label_key=None, 201 patch_shape=patch_shape, 202 is_seg_dataset=False, 203 **kwargs 204 )
Get the OCTA-500 dataset for retinal vessel and FAZ segmentation in OCTA en-face projections.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- subset: The choice of field-of-view subset. Either '3M' or '6M'.
- label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 'vein' or 'faz'.
- projection: The choice of 2D en-face projection map used as the raw input.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
207def get_octa500_loader( 208 path: Union[os.PathLike, str], 209 batch_size: int, 210 patch_shape: Tuple[int, int], 211 subset: Literal["3M", "6M"], 212 label_type: Literal["large_vessel", "capillary", "artery", "vein", "faz"] = "large_vessel", 213 projection: Literal[ 214 "oct_full", "oct_ilm_opl", "oct_opl_bm", "octa_full", "octa_ilm_opl", "octa_opl_bm" 215 ] = "octa_full", 216 resize_inputs: bool = False, 217 download: bool = False, 218 **kwargs 219) -> DataLoader: 220 """Get the OCTA-500 dataloader for retinal vessel and FAZ segmentation in OCTA en-face projections. 221 222 Args: 223 path: Filepath to a folder where the data is downloaded for further processing. 224 batch_size: The batch size for training. 225 patch_shape: The patch shape to use for training. 226 subset: The choice of field-of-view subset. Either '3M' or '6M'. 227 label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 228 'vein' or 'faz'. 229 projection: The choice of 2D en-face projection map used as the raw input. 230 resize_inputs: Whether to resize the inputs to the expected patch shape. 231 download: Whether to download the data if it is not present. 232 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 233 234 Returns: 235 The DataLoader. 236 """ 237 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 238 dataset = get_octa500_dataset( 239 path, patch_shape, subset, label_type, projection, resize_inputs, download, **ds_kwargs 240 ) 241 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the OCTA-500 dataloader for retinal vessel and FAZ segmentation in OCTA en-face projections.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- subset: The choice of field-of-view subset. Either '3M' or '6M'.
- label_type: The choice of segmentation label. One of 'large_vessel', 'capillary', 'artery', 'vein' or 'faz'.
- projection: The choice of 2D en-face projection map used as the raw input.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.