torch_em.data.datasets.medical.derma_octa
The DERMA-OCTA dataset contains annotations for dermatological vessel segmentation in optical coherence tomography angiography (OCTA) images, comprising 330 volumetric acquisitions from 74 subjects (chronic venous disease, healthy skin and skin lesions).
NOTE: The dataset abstract mentions that segmentation labels shown in the paper's figures were generated with a U-Net. The archive itself, however, ships a dedicated 'Manual segmentations' folder (both for the 2D en-face projections and the 3D volumes), which are the real reference annotations created in Amira and manually refined/verified by four annotators plus a technician and a clinician. This loader only uses these manual segmentations, never any U-Net-generated files.
The raw acquisitions are only distributed after one of five preprocessing pipelines has been applied to them (the archive does not ship the untouched raw scans separately). This loader uses the 'Norm' (min-max normalized) version, paired 1:1 with the manual segmentations by filename.
Each case provides three mask variants: the full-depth projection ('all'), the superficial vascular plexus ('sup') and the deep vascular plexus ('deep'), selected with the 'plexus' argument (2D only; the 3D volumes only ship the full-depth variant).
NOTE: A single case in the 3D archive ('13801') ships a label volume with one extra frame relative to its raw volume; this loader drops that mismatched pair defensively.
The data is located at https://doi.org/10.5281/zenodo.15088516, released under a CC-BY-NC-4.0 license. This dataset is from the publication https://doi.org/10.1038/s41597-025-05763-6. Please cite it if you use this dataset for your research.
1"""The DERMA-OCTA dataset contains annotations for dermatological vessel segmentation in optical 2coherence tomography angiography (OCTA) images, comprising 330 volumetric acquisitions from 74 3subjects (chronic venous disease, healthy skin and skin lesions). 4 5NOTE: The dataset abstract mentions that segmentation labels shown in the paper's figures were 6generated with a U-Net. The archive itself, however, ships a dedicated 'Manual segmentations' folder 7(both for the 2D en-face projections and the 3D volumes), which are the real reference annotations 8created in Amira and manually refined/verified by four annotators plus a technician and a clinician. 9This loader only uses these manual segmentations, never any U-Net-generated files. 10 11The raw acquisitions are only distributed after one of five preprocessing pipelines has been applied 12to them (the archive does not ship the untouched raw scans separately). This loader uses the 'Norm' 13(min-max normalized) version, paired 1:1 with the manual segmentations by filename. 14 15Each case provides three mask variants: the full-depth projection ('all'), the superficial vascular 16plexus ('sup') and the deep vascular plexus ('deep'), selected with the 'plexus' argument (2D only; 17the 3D volumes only ship the full-depth variant). 18 19NOTE: A single case in the 3D archive ('13801') ships a label volume with one extra frame relative to 20its raw volume; this loader drops that mismatched pair defensively. 21 22The data is located at https://doi.org/10.5281/zenodo.15088516, released under a CC-BY-NC-4.0 license. 23This dataset is from the publication https://doi.org/10.1038/s41597-025-05763-6. 24Please cite it if you use this dataset for your research. 25""" 26 27import os 28import re 29from glob import glob 30from natsort import natsorted 31from typing import Union, Tuple, Literal, List 32 33from torch.utils.data import Dataset, DataLoader 34 35import torch_em 36 37from .. import util 38 39 40URLS = { 41 "2d": { 42 "raw": "https://zenodo.org/api/records/15088516/files/Norm%20-%202D.zip/content", 43 "labels": "https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%202D.zip/content", 44 }, 45 "3d": { 46 "raw": "https://zenodo.org/api/records/15088516/files/Norm%20-%203D.zip/content", 47 "labels": "https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%203D.zip/content", 48 }, 49} 50 51CHECKSUMS = { 52 "2d": { 53 "raw": "f481652c2753d0d23f56cba1c52cc980c9075dff2a144805ff30f7003f6e1c62", 54 "labels": "b4971dba5dc80d0322396300a9f40cc554205bd68b5734b2f5f4c449baa7b1b4", 55 }, 56 "3d": { 57 "raw": "5c226bc7dd2f625463473f8cc6b448719f7efa05909b9fc664acc17ce9e218bb", 58 "labels": "6155e49943dbe591e3b112e0a2db643a622f13a2fd4857d559b699be9724ac1a", 59 }, 60} 61 62DIR_NAMES = {"2d": ("Norm - 2D", "Manual segmentations - 2D"), "3d": ("Norm - 3D", "Manual segmentations - 3D")} 63 64 65def get_derma_octa_data(path: Union[os.PathLike, str], dim: Literal["2d", "3d"] = "2d", download: bool = False) -> str: 66 """Download the DERMA-OCTA dataset. 67 68 Args: 69 path: Filepath to a folder where the data is downloaded for further processing. 70 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 71 download: Whether to download the data if it is not present. 72 73 Returns: 74 Filepath where the data is downloaded. 75 """ 76 if dim not in URLS: 77 raise ValueError(f"'{dim}' is not a valid choice of dimensionality. Choose one of {list(URLS.keys())}.") 78 79 os.makedirs(path, exist_ok=True) 80 81 raw_dirname, label_dirname = DIR_NAMES[dim] 82 raw_dir = os.path.join(path, raw_dirname) 83 label_dir = os.path.join(path, label_dirname) 84 85 if not os.path.exists(raw_dir): 86 raw_zip = os.path.join(path, f"Norm_{dim}.zip") 87 util.download_source(path=raw_zip, url=URLS[dim]["raw"], download=download, checksum=CHECKSUMS[dim]["raw"]) 88 util.unzip(zip_path=raw_zip, dst=path) 89 90 if not os.path.exists(label_dir): 91 label_zip = os.path.join(path, f"Manual_segmentations_{dim}.zip") 92 util.download_source( 93 path=label_zip, url=URLS[dim]["labels"], download=download, checksum=CHECKSUMS[dim]["labels"] 94 ) 95 util.unzip(zip_path=label_zip, dst=path) 96 97 return path 98 99 100def get_derma_octa_paths( 101 path: Union[os.PathLike, str], 102 dim: Literal["2d", "3d"] = "2d", 103 plexus: Literal["all", "sup", "deep"] = "all", 104 download: bool = False, 105) -> Tuple[List[str], List[str]]: 106 """Get paths to the DERMA-OCTA data. 107 108 Args: 109 path: Filepath to a folder where the data is downloaded for further processing. 110 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 111 plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or 112 'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant. 113 download: Whether to download the data if it is not present. 114 115 Returns: 116 List of filepaths for the image data. 117 List of filepaths for the label data. 118 """ 119 data_dir = get_derma_octa_data(path, dim, download) 120 raw_dirname, label_dirname = DIR_NAMES[dim] 121 122 if dim == "2d": 123 label_paths = natsorted(glob(os.path.join(data_dir, label_dirname, "*", f"*_{plexus}.png"))) 124 raw_paths = [ 125 os.path.join(data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), os.path.basename(p)) 126 for p in label_paths 127 ] 128 else: 129 label_paths = natsorted( 130 p for p in glob(os.path.join(data_dir, label_dirname, "*", "*.tif")) if re.match(r"^\d+\.tif$", os.path.basename(p)) # noqa 131 ) 132 raw_paths = [ 133 os.path.join( 134 data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), 135 f"norm_{os.path.splitext(os.path.basename(p))[0]}.tiff" 136 ) 137 for p in label_paths 138 ] 139 140 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 141 assert all(os.path.exists(p) for p in raw_paths) 142 143 if dim == "3d": 144 # A single case in the archive ('13801') ships a label volume with one extra frame relative 145 # to its raw volume (90 vs. 91 slices). Drop any such mismatched pairs defensively. 146 import tifffile 147 148 def _shape(p): 149 pages = tifffile.TiffFile(p).pages 150 return (len(pages),) + pages[0].shape 151 152 filtered = [(r, lb) for r, lb in zip(raw_paths, label_paths) if _shape(r) == _shape(lb)] 153 raw_paths, label_paths = map(list, zip(*filtered)) 154 155 return raw_paths, label_paths 156 157 158def get_derma_octa_dataset( 159 path: Union[os.PathLike, str], 160 patch_shape: Tuple[int, ...], 161 dim: Literal["2d", "3d"] = "2d", 162 plexus: Literal["all", "sup", "deep"] = "all", 163 resize_inputs: bool = False, 164 download: bool = False, 165 **kwargs 166) -> Dataset: 167 """Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images. 168 169 Args: 170 path: Filepath to a folder where the data is downloaded for further processing. 171 patch_shape: The patch shape to use for training. 172 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 173 plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details. 174 resize_inputs: Whether to resize the inputs to the patch shape. 175 download: Whether to download the data if it is not present. 176 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 177 178 Returns: 179 The segmentation dataset. 180 """ 181 raw_paths, label_paths = get_derma_octa_paths(path, dim, plexus, download) 182 183 if resize_inputs: 184 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": dim == "2d"} 185 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 186 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 187 ) 188 189 return torch_em.default_segmentation_dataset( 190 raw_paths=raw_paths, 191 raw_key=None, 192 label_paths=label_paths, 193 label_key=None, 194 is_seg_dataset=(dim == "3d"), 195 patch_shape=patch_shape, 196 ndim=2 if dim == "2d" else 3, 197 **kwargs 198 ) 199 200 201def get_derma_octa_loader( 202 path: Union[os.PathLike, str], 203 batch_size: int, 204 patch_shape: Tuple[int, ...], 205 dim: Literal["2d", "3d"] = "2d", 206 plexus: Literal["all", "sup", "deep"] = "all", 207 resize_inputs: bool = False, 208 download: bool = False, 209 **kwargs 210) -> DataLoader: 211 """Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images. 212 213 Args: 214 path: Filepath to a folder where the data is downloaded for further processing. 215 batch_size: The batch size for training. 216 patch_shape: The patch shape to use for training. 217 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 218 plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details. 219 resize_inputs: Whether to resize the inputs to the patch shape. 220 download: Whether to download the data if it is not present. 221 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 222 223 Returns: 224 The DataLoader. 225 """ 226 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 227 dataset = get_derma_octa_dataset(path, patch_shape, dim, plexus, resize_inputs, download, **ds_kwargs) 228 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
66def get_derma_octa_data(path: Union[os.PathLike, str], dim: Literal["2d", "3d"] = "2d", download: bool = False) -> str: 67 """Download the DERMA-OCTA dataset. 68 69 Args: 70 path: Filepath to a folder where the data is downloaded for further processing. 71 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 72 download: Whether to download the data if it is not present. 73 74 Returns: 75 Filepath where the data is downloaded. 76 """ 77 if dim not in URLS: 78 raise ValueError(f"'{dim}' is not a valid choice of dimensionality. Choose one of {list(URLS.keys())}.") 79 80 os.makedirs(path, exist_ok=True) 81 82 raw_dirname, label_dirname = DIR_NAMES[dim] 83 raw_dir = os.path.join(path, raw_dirname) 84 label_dir = os.path.join(path, label_dirname) 85 86 if not os.path.exists(raw_dir): 87 raw_zip = os.path.join(path, f"Norm_{dim}.zip") 88 util.download_source(path=raw_zip, url=URLS[dim]["raw"], download=download, checksum=CHECKSUMS[dim]["raw"]) 89 util.unzip(zip_path=raw_zip, dst=path) 90 91 if not os.path.exists(label_dir): 92 label_zip = os.path.join(path, f"Manual_segmentations_{dim}.zip") 93 util.download_source( 94 path=label_zip, url=URLS[dim]["labels"], download=download, checksum=CHECKSUMS[dim]["labels"] 95 ) 96 util.unzip(zip_path=label_zip, dst=path) 97 98 return path
Download the DERMA-OCTA dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is downloaded.
101def get_derma_octa_paths( 102 path: Union[os.PathLike, str], 103 dim: Literal["2d", "3d"] = "2d", 104 plexus: Literal["all", "sup", "deep"] = "all", 105 download: bool = False, 106) -> Tuple[List[str], List[str]]: 107 """Get paths to the DERMA-OCTA data. 108 109 Args: 110 path: Filepath to a folder where the data is downloaded for further processing. 111 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 112 plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or 113 'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant. 114 download: Whether to download the data if it is not present. 115 116 Returns: 117 List of filepaths for the image data. 118 List of filepaths for the label data. 119 """ 120 data_dir = get_derma_octa_data(path, dim, download) 121 raw_dirname, label_dirname = DIR_NAMES[dim] 122 123 if dim == "2d": 124 label_paths = natsorted(glob(os.path.join(data_dir, label_dirname, "*", f"*_{plexus}.png"))) 125 raw_paths = [ 126 os.path.join(data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), os.path.basename(p)) 127 for p in label_paths 128 ] 129 else: 130 label_paths = natsorted( 131 p for p in glob(os.path.join(data_dir, label_dirname, "*", "*.tif")) if re.match(r"^\d+\.tif$", os.path.basename(p)) # noqa 132 ) 133 raw_paths = [ 134 os.path.join( 135 data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), 136 f"norm_{os.path.splitext(os.path.basename(p))[0]}.tiff" 137 ) 138 for p in label_paths 139 ] 140 141 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 142 assert all(os.path.exists(p) for p in raw_paths) 143 144 if dim == "3d": 145 # A single case in the archive ('13801') ships a label volume with one extra frame relative 146 # to its raw volume (90 vs. 91 slices). Drop any such mismatched pairs defensively. 147 import tifffile 148 149 def _shape(p): 150 pages = tifffile.TiffFile(p).pages 151 return (len(pages),) + pages[0].shape 152 153 filtered = [(r, lb) for r, lb in zip(raw_paths, label_paths) if _shape(r) == _shape(lb)] 154 raw_paths, label_paths = map(list, zip(*filtered)) 155 156 return raw_paths, label_paths
Get paths to the DERMA-OCTA data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
- plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or 'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
159def get_derma_octa_dataset( 160 path: Union[os.PathLike, str], 161 patch_shape: Tuple[int, ...], 162 dim: Literal["2d", "3d"] = "2d", 163 plexus: Literal["all", "sup", "deep"] = "all", 164 resize_inputs: bool = False, 165 download: bool = False, 166 **kwargs 167) -> Dataset: 168 """Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images. 169 170 Args: 171 path: Filepath to a folder where the data is downloaded for further processing. 172 patch_shape: The patch shape to use for training. 173 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 174 plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details. 175 resize_inputs: Whether to resize the inputs to the patch shape. 176 download: Whether to download the data if it is not present. 177 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 178 179 Returns: 180 The segmentation dataset. 181 """ 182 raw_paths, label_paths = get_derma_octa_paths(path, dim, plexus, download) 183 184 if resize_inputs: 185 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": dim == "2d"} 186 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 187 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 188 ) 189 190 return torch_em.default_segmentation_dataset( 191 raw_paths=raw_paths, 192 raw_key=None, 193 label_paths=label_paths, 194 label_key=None, 195 is_seg_dataset=(dim == "3d"), 196 patch_shape=patch_shape, 197 ndim=2 if dim == "2d" else 3, 198 **kwargs 199 )
Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
- plexus: The choice of vascular plexus mask. See
get_derma_octa_pathsfor details. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
202def get_derma_octa_loader( 203 path: Union[os.PathLike, str], 204 batch_size: int, 205 patch_shape: Tuple[int, ...], 206 dim: Literal["2d", "3d"] = "2d", 207 plexus: Literal["all", "sup", "deep"] = "all", 208 resize_inputs: bool = False, 209 download: bool = False, 210 **kwargs 211) -> DataLoader: 212 """Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images. 213 214 Args: 215 path: Filepath to a folder where the data is downloaded for further processing. 216 batch_size: The batch size for training. 217 patch_shape: The patch shape to use for training. 218 dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes). 219 plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details. 220 resize_inputs: Whether to resize the inputs to the patch shape. 221 download: Whether to download the data if it is not present. 222 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 223 224 Returns: 225 The DataLoader. 226 """ 227 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 228 dataset = get_derma_octa_dataset(path, patch_shape, dim, plexus, resize_inputs, download, **ds_kwargs) 229 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
- plexus: The choice of vascular plexus mask. See
get_derma_octa_pathsfor details. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.