torch_em.data.datasets.medical.derma_octa

The DERMA-OCTA dataset contains annotations for dermatological vessel segmentation in optical coherence tomography angiography (OCTA) images, comprising 330 volumetric acquisitions from 74 subjects (chronic venous disease, healthy skin and skin lesions).

NOTE: The dataset abstract mentions that segmentation labels shown in the paper's figures were generated with a U-Net. The archive itself, however, ships a dedicated 'Manual segmentations' folder (both for the 2D en-face projections and the 3D volumes), which are the real reference annotations created in Amira and manually refined/verified by four annotators plus a technician and a clinician. This loader only uses these manual segmentations, never any U-Net-generated files.

The raw acquisitions are only distributed after one of five preprocessing pipelines has been applied to them (the archive does not ship the untouched raw scans separately). This loader uses the 'Norm' (min-max normalized) version, paired 1:1 with the manual segmentations by filename.

Each case provides three mask variants: the full-depth projection ('all'), the superficial vascular plexus ('sup') and the deep vascular plexus ('deep'), selected with the 'plexus' argument (2D only; the 3D volumes only ship the full-depth variant).

NOTE: A single case in the 3D archive ('13801') ships a label volume with one extra frame relative to its raw volume; this loader drops that mismatched pair defensively.

The data is located at https://doi.org/10.5281/zenodo.15088516, released under a CC-BY-NC-4.0 license. This dataset is from the publication https://doi.org/10.1038/s41597-025-05763-6. Please cite it if you use this dataset for your research.

  1"""The DERMA-OCTA dataset contains annotations for dermatological vessel segmentation in optical
  2coherence tomography angiography (OCTA) images, comprising 330 volumetric acquisitions from 74
  3subjects (chronic venous disease, healthy skin and skin lesions).
  4
  5NOTE: The dataset abstract mentions that segmentation labels shown in the paper's figures were
  6generated with a U-Net. The archive itself, however, ships a dedicated 'Manual segmentations' folder
  7(both for the 2D en-face projections and the 3D volumes), which are the real reference annotations
  8created in Amira and manually refined/verified by four annotators plus a technician and a clinician.
  9This loader only uses these manual segmentations, never any U-Net-generated files.
 10
 11The raw acquisitions are only distributed after one of five preprocessing pipelines has been applied
 12to them (the archive does not ship the untouched raw scans separately). This loader uses the 'Norm'
 13(min-max normalized) version, paired 1:1 with the manual segmentations by filename.
 14
 15Each case provides three mask variants: the full-depth projection ('all'), the superficial vascular
 16plexus ('sup') and the deep vascular plexus ('deep'), selected with the 'plexus' argument (2D only;
 17the 3D volumes only ship the full-depth variant).
 18
 19NOTE: A single case in the 3D archive ('13801') ships a label volume with one extra frame relative to
 20its raw volume; this loader drops that mismatched pair defensively.
 21
 22The data is located at https://doi.org/10.5281/zenodo.15088516, released under a CC-BY-NC-4.0 license.
 23This dataset is from the publication https://doi.org/10.1038/s41597-025-05763-6.
 24Please cite it if you use this dataset for your research.
 25"""
 26
 27import os
 28import re
 29from glob import glob
 30from natsort import natsorted
 31from typing import Union, Tuple, Literal, List
 32
 33from torch.utils.data import Dataset, DataLoader
 34
 35import torch_em
 36
 37from .. import util
 38
 39
 40URLS = {
 41    "2d": {
 42        "raw": "https://zenodo.org/api/records/15088516/files/Norm%20-%202D.zip/content",
 43        "labels": "https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%202D.zip/content",
 44    },
 45    "3d": {
 46        "raw": "https://zenodo.org/api/records/15088516/files/Norm%20-%203D.zip/content",
 47        "labels": "https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%203D.zip/content",
 48    },
 49}
 50
 51CHECKSUMS = {
 52    "2d": {
 53        "raw": "f481652c2753d0d23f56cba1c52cc980c9075dff2a144805ff30f7003f6e1c62",
 54        "labels": "b4971dba5dc80d0322396300a9f40cc554205bd68b5734b2f5f4c449baa7b1b4",
 55    },
 56    "3d": {
 57        "raw": "5c226bc7dd2f625463473f8cc6b448719f7efa05909b9fc664acc17ce9e218bb",
 58        "labels": "6155e49943dbe591e3b112e0a2db643a622f13a2fd4857d559b699be9724ac1a",
 59    },
 60}
 61
 62DIR_NAMES = {"2d": ("Norm - 2D", "Manual segmentations - 2D"), "3d": ("Norm - 3D", "Manual segmentations - 3D")}
 63
 64
 65def get_derma_octa_data(path: Union[os.PathLike, str], dim: Literal["2d", "3d"] = "2d", download: bool = False) -> str:
 66    """Download the DERMA-OCTA dataset.
 67
 68    Args:
 69        path: Filepath to a folder where the data is downloaded for further processing.
 70        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
 71        download: Whether to download the data if it is not present.
 72
 73    Returns:
 74        Filepath where the data is downloaded.
 75    """
 76    if dim not in URLS:
 77        raise ValueError(f"'{dim}' is not a valid choice of dimensionality. Choose one of {list(URLS.keys())}.")
 78
 79    os.makedirs(path, exist_ok=True)
 80
 81    raw_dirname, label_dirname = DIR_NAMES[dim]
 82    raw_dir = os.path.join(path, raw_dirname)
 83    label_dir = os.path.join(path, label_dirname)
 84
 85    if not os.path.exists(raw_dir):
 86        raw_zip = os.path.join(path, f"Norm_{dim}.zip")
 87        util.download_source(path=raw_zip, url=URLS[dim]["raw"], download=download, checksum=CHECKSUMS[dim]["raw"])
 88        util.unzip(zip_path=raw_zip, dst=path)
 89
 90    if not os.path.exists(label_dir):
 91        label_zip = os.path.join(path, f"Manual_segmentations_{dim}.zip")
 92        util.download_source(
 93            path=label_zip, url=URLS[dim]["labels"], download=download, checksum=CHECKSUMS[dim]["labels"]
 94        )
 95        util.unzip(zip_path=label_zip, dst=path)
 96
 97    return path
 98
 99
100def get_derma_octa_paths(
101    path: Union[os.PathLike, str],
102    dim: Literal["2d", "3d"] = "2d",
103    plexus: Literal["all", "sup", "deep"] = "all",
104    download: bool = False,
105) -> Tuple[List[str], List[str]]:
106    """Get paths to the DERMA-OCTA data.
107
108    Args:
109        path: Filepath to a folder where the data is downloaded for further processing.
110        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
111        plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or
112            'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant.
113        download: Whether to download the data if it is not present.
114
115    Returns:
116        List of filepaths for the image data.
117        List of filepaths for the label data.
118    """
119    data_dir = get_derma_octa_data(path, dim, download)
120    raw_dirname, label_dirname = DIR_NAMES[dim]
121
122    if dim == "2d":
123        label_paths = natsorted(glob(os.path.join(data_dir, label_dirname, "*", f"*_{plexus}.png")))
124        raw_paths = [
125            os.path.join(data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), os.path.basename(p))
126            for p in label_paths
127        ]
128    else:
129        label_paths = natsorted(
130            p for p in glob(os.path.join(data_dir, label_dirname, "*", "*.tif")) if re.match(r"^\d+\.tif$", os.path.basename(p))  # noqa
131        )
132        raw_paths = [
133            os.path.join(
134                data_dir, raw_dirname, os.path.basename(os.path.dirname(p)),
135                f"norm_{os.path.splitext(os.path.basename(p))[0]}.tiff"
136            )
137            for p in label_paths
138        ]
139
140    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
141    assert all(os.path.exists(p) for p in raw_paths)
142
143    if dim == "3d":
144        # A single case in the archive ('13801') ships a label volume with one extra frame relative
145        # to its raw volume (90 vs. 91 slices). Drop any such mismatched pairs defensively.
146        import tifffile
147
148        def _shape(p):
149            pages = tifffile.TiffFile(p).pages
150            return (len(pages),) + pages[0].shape
151
152        filtered = [(r, lb) for r, lb in zip(raw_paths, label_paths) if _shape(r) == _shape(lb)]
153        raw_paths, label_paths = map(list, zip(*filtered))
154
155    return raw_paths, label_paths
156
157
158def get_derma_octa_dataset(
159    path: Union[os.PathLike, str],
160    patch_shape: Tuple[int, ...],
161    dim: Literal["2d", "3d"] = "2d",
162    plexus: Literal["all", "sup", "deep"] = "all",
163    resize_inputs: bool = False,
164    download: bool = False,
165    **kwargs
166) -> Dataset:
167    """Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images.
168
169    Args:
170        path: Filepath to a folder where the data is downloaded for further processing.
171        patch_shape: The patch shape to use for training.
172        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
173        plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details.
174        resize_inputs: Whether to resize the inputs to the patch shape.
175        download: Whether to download the data if it is not present.
176        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
177
178    Returns:
179        The segmentation dataset.
180    """
181    raw_paths, label_paths = get_derma_octa_paths(path, dim, plexus, download)
182
183    if resize_inputs:
184        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": dim == "2d"}
185        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
186            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
187        )
188
189    return torch_em.default_segmentation_dataset(
190        raw_paths=raw_paths,
191        raw_key=None,
192        label_paths=label_paths,
193        label_key=None,
194        is_seg_dataset=(dim == "3d"),
195        patch_shape=patch_shape,
196        ndim=2 if dim == "2d" else 3,
197        **kwargs
198    )
199
200
201def get_derma_octa_loader(
202    path: Union[os.PathLike, str],
203    batch_size: int,
204    patch_shape: Tuple[int, ...],
205    dim: Literal["2d", "3d"] = "2d",
206    plexus: Literal["all", "sup", "deep"] = "all",
207    resize_inputs: bool = False,
208    download: bool = False,
209    **kwargs
210) -> DataLoader:
211    """Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images.
212
213    Args:
214        path: Filepath to a folder where the data is downloaded for further processing.
215        batch_size: The batch size for training.
216        patch_shape: The patch shape to use for training.
217        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
218        plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details.
219        resize_inputs: Whether to resize the inputs to the patch shape.
220        download: Whether to download the data if it is not present.
221        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
222
223    Returns:
224        The DataLoader.
225    """
226    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
227    dataset = get_derma_octa_dataset(path, patch_shape, dim, plexus, resize_inputs, download, **ds_kwargs)
228    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URLS = {'2d': {'raw': 'https://zenodo.org/api/records/15088516/files/Norm%20-%202D.zip/content', 'labels': 'https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%202D.zip/content'}, '3d': {'raw': 'https://zenodo.org/api/records/15088516/files/Norm%20-%203D.zip/content', 'labels': 'https://zenodo.org/api/records/15088516/files/Manual%20segmentations%20-%203D.zip/content'}}
CHECKSUMS = {'2d': {'raw': 'f481652c2753d0d23f56cba1c52cc980c9075dff2a144805ff30f7003f6e1c62', 'labels': 'b4971dba5dc80d0322396300a9f40cc554205bd68b5734b2f5f4c449baa7b1b4'}, '3d': {'raw': '5c226bc7dd2f625463473f8cc6b448719f7efa05909b9fc664acc17ce9e218bb', 'labels': '6155e49943dbe591e3b112e0a2db643a622f13a2fd4857d559b699be9724ac1a'}}
DIR_NAMES = {'2d': ('Norm - 2D', 'Manual segmentations - 2D'), '3d': ('Norm - 3D', 'Manual segmentations - 3D')}
def get_derma_octa_data( path: Union[os.PathLike, str], dim: Literal['2d', '3d'] = '2d', download: bool = False) -> str:
66def get_derma_octa_data(path: Union[os.PathLike, str], dim: Literal["2d", "3d"] = "2d", download: bool = False) -> str:
67    """Download the DERMA-OCTA dataset.
68
69    Args:
70        path: Filepath to a folder where the data is downloaded for further processing.
71        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
72        download: Whether to download the data if it is not present.
73
74    Returns:
75        Filepath where the data is downloaded.
76    """
77    if dim not in URLS:
78        raise ValueError(f"'{dim}' is not a valid choice of dimensionality. Choose one of {list(URLS.keys())}.")
79
80    os.makedirs(path, exist_ok=True)
81
82    raw_dirname, label_dirname = DIR_NAMES[dim]
83    raw_dir = os.path.join(path, raw_dirname)
84    label_dir = os.path.join(path, label_dirname)
85
86    if not os.path.exists(raw_dir):
87        raw_zip = os.path.join(path, f"Norm_{dim}.zip")
88        util.download_source(path=raw_zip, url=URLS[dim]["raw"], download=download, checksum=CHECKSUMS[dim]["raw"])
89        util.unzip(zip_path=raw_zip, dst=path)
90
91    if not os.path.exists(label_dir):
92        label_zip = os.path.join(path, f"Manual_segmentations_{dim}.zip")
93        util.download_source(
94            path=label_zip, url=URLS[dim]["labels"], download=download, checksum=CHECKSUMS[dim]["labels"]
95        )
96        util.unzip(zip_path=label_zip, dst=path)
97
98    return path

Download the DERMA-OCTA dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_derma_octa_paths( path: Union[os.PathLike, str], dim: Literal['2d', '3d'] = '2d', plexus: Literal['all', 'sup', 'deep'] = 'all', download: bool = False) -> Tuple[List[str], List[str]]:
101def get_derma_octa_paths(
102    path: Union[os.PathLike, str],
103    dim: Literal["2d", "3d"] = "2d",
104    plexus: Literal["all", "sup", "deep"] = "all",
105    download: bool = False,
106) -> Tuple[List[str], List[str]]:
107    """Get paths to the DERMA-OCTA data.
108
109    Args:
110        path: Filepath to a folder where the data is downloaded for further processing.
111        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
112        plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or
113            'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant.
114        download: Whether to download the data if it is not present.
115
116    Returns:
117        List of filepaths for the image data.
118        List of filepaths for the label data.
119    """
120    data_dir = get_derma_octa_data(path, dim, download)
121    raw_dirname, label_dirname = DIR_NAMES[dim]
122
123    if dim == "2d":
124        label_paths = natsorted(glob(os.path.join(data_dir, label_dirname, "*", f"*_{plexus}.png")))
125        raw_paths = [
126            os.path.join(data_dir, raw_dirname, os.path.basename(os.path.dirname(p)), os.path.basename(p))
127            for p in label_paths
128        ]
129    else:
130        label_paths = natsorted(
131            p for p in glob(os.path.join(data_dir, label_dirname, "*", "*.tif")) if re.match(r"^\d+\.tif$", os.path.basename(p))  # noqa
132        )
133        raw_paths = [
134            os.path.join(
135                data_dir, raw_dirname, os.path.basename(os.path.dirname(p)),
136                f"norm_{os.path.splitext(os.path.basename(p))[0]}.tiff"
137            )
138            for p in label_paths
139        ]
140
141    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
142    assert all(os.path.exists(p) for p in raw_paths)
143
144    if dim == "3d":
145        # A single case in the archive ('13801') ships a label volume with one extra frame relative
146        # to its raw volume (90 vs. 91 slices). Drop any such mismatched pairs defensively.
147        import tifffile
148
149        def _shape(p):
150            pages = tifffile.TiffFile(p).pages
151            return (len(pages),) + pages[0].shape
152
153        filtered = [(r, lb) for r, lb in zip(raw_paths, label_paths) if _shape(r) == _shape(lb)]
154        raw_paths, label_paths = map(list, zip(*filtered))
155
156    return raw_paths, label_paths

Get paths to the DERMA-OCTA data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
  • plexus: The choice of vascular plexus mask. One of 'all' (full-depth), 'sup' (superficial) or 'deep' (deep). Only relevant for 'dim=2d': the 3D volumes only ship the 'all' variant.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_derma_octa_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], dim: Literal['2d', '3d'] = '2d', plexus: Literal['all', 'sup', 'deep'] = 'all', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
159def get_derma_octa_dataset(
160    path: Union[os.PathLike, str],
161    patch_shape: Tuple[int, ...],
162    dim: Literal["2d", "3d"] = "2d",
163    plexus: Literal["all", "sup", "deep"] = "all",
164    resize_inputs: bool = False,
165    download: bool = False,
166    **kwargs
167) -> Dataset:
168    """Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images.
169
170    Args:
171        path: Filepath to a folder where the data is downloaded for further processing.
172        patch_shape: The patch shape to use for training.
173        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
174        plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details.
175        resize_inputs: Whether to resize the inputs to the patch shape.
176        download: Whether to download the data if it is not present.
177        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
178
179    Returns:
180        The segmentation dataset.
181    """
182    raw_paths, label_paths = get_derma_octa_paths(path, dim, plexus, download)
183
184    if resize_inputs:
185        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": dim == "2d"}
186        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
187            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
188        )
189
190    return torch_em.default_segmentation_dataset(
191        raw_paths=raw_paths,
192        raw_key=None,
193        label_paths=label_paths,
194        label_key=None,
195        is_seg_dataset=(dim == "3d"),
196        patch_shape=patch_shape,
197        ndim=2 if dim == "2d" else 3,
198        **kwargs
199    )

Get the DERMA-OCTA dataset for dermatological vessel segmentation in OCTA images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
  • plexus: The choice of vascular plexus mask. See get_derma_octa_paths for details.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_derma_octa_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], dim: Literal['2d', '3d'] = '2d', plexus: Literal['all', 'sup', 'deep'] = 'all', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
202def get_derma_octa_loader(
203    path: Union[os.PathLike, str],
204    batch_size: int,
205    patch_shape: Tuple[int, ...],
206    dim: Literal["2d", "3d"] = "2d",
207    plexus: Literal["all", "sup", "deep"] = "all",
208    resize_inputs: bool = False,
209    download: bool = False,
210    **kwargs
211) -> DataLoader:
212    """Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images.
213
214    Args:
215        path: Filepath to a folder where the data is downloaded for further processing.
216        batch_size: The batch size for training.
217        patch_shape: The patch shape to use for training.
218        dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
219        plexus: The choice of vascular plexus mask. See `get_derma_octa_paths` for details.
220        resize_inputs: Whether to resize the inputs to the patch shape.
221        download: Whether to download the data if it is not present.
222        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
223
224    Returns:
225        The DataLoader.
226    """
227    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
228    dataset = get_derma_octa_dataset(path, patch_shape, dim, plexus, resize_inputs, download, **ds_kwargs)
229    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the DERMA-OCTA dataloader for dermatological vessel segmentation in OCTA images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • dim: The choice of data dimensionality. Either '2d' (en-face projections) or '3d' (volumes).
  • plexus: The choice of vascular plexus mask. See get_derma_octa_paths for details.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.