torch_em.data.datasets.light_microscopy.air_leish

The AIR-LEISH dataset contains annotations for the segmentation of Leishmania amastigotes, host cells (macrophages) and nuclei in microscopy images of Giemsa-stained Leishmania-infected macrophages.

The dataset consists of 180 RGB images of 1844 x 2709 pixels, in two sets of 90 images (Set1, file names '20250328_CC*', and Set2, file names '20250203_CF*'). Every image is annotated with three classes, 'AM' (amastigotes), 'HC' (host cells) and 'NU' (nuclei), as COCO polygons (one polygon per instance) and, with one exception, as a pixel-wise class mask (0: background, 1: amastigote, 2: host cell, 3: nucleus). In the class masks the nuclei are cut out of the host cells (the polygon of a host cell covers its nucleus).

The target argument selects the labels:

  • 'semantic': the class masks shipped with the dataset. Image '20250328_CCimage49' of Set1 has no class mask and is not part of this target (179 images).
  • 'amastigotes', 'host_cells' or 'nuclei': instance labels rasterized from the COCO polygons of this class (all 180 images, one id per instance in the order of the annotation file, later instances overwrite earlier ones where they overlap, which is rare). 26 images do not contain any amastigote and have an empty label.

NOTE: The images are stored as RGBA and are converted to RGB, and the labels of the instance targets are rasterized once when the paths are first requested. 54 file names in the annotations of Set2 have a doubled '.png' extension, which is handled by this module.

The data is located at https://doi.org/10.5281/zenodo.17384855 and released under a CC-BY-4.0 license. The Zenodo record also contains a second, smaller archive ('AIR-Leish_dataset.zip'), which holds 67 of the same images and is not used here.

Please cite the publication associated with the Zenodo record if you use this dataset in your research.

  1"""The AIR-LEISH dataset contains annotations for the segmentation of Leishmania amastigotes, host cells
  2(macrophages) and nuclei in microscopy images of Giemsa-stained Leishmania-infected macrophages.
  3
  4The dataset consists of 180 RGB images of 1844 x 2709 pixels, in two sets of 90 images (`Set1`, file names
  5'20250328_CC*', and `Set2`, file names '20250203_CF*'). Every image is annotated with three classes,
  6'AM' (amastigotes), 'HC' (host cells) and 'NU' (nuclei), as COCO polygons (one polygon per instance) and,
  7with one exception, as a pixel-wise class mask (0: background, 1: amastigote, 2: host cell, 3: nucleus).
  8In the class masks the nuclei are cut out of the host cells (the polygon of a host cell covers its nucleus).
  9
 10The `target` argument selects the labels:
 11- 'semantic': the class masks shipped with the dataset. Image '20250328_CCimage49' of `Set1` has no class mask
 12  and is not part of this target (179 images).
 13- 'amastigotes', 'host_cells' or 'nuclei': instance labels rasterized from the COCO polygons of this class
 14  (all 180 images, one id per instance in the order of the annotation file, later instances overwrite earlier
 15  ones where they overlap, which is rare). 26 images do not contain any amastigote and have an empty label.
 16
 17NOTE: The images are stored as RGBA and are converted to RGB, and the labels of the instance targets are
 18rasterized once when the paths are first requested. 54 file names in the annotations of `Set2` have a doubled
 19'.png' extension, which is handled by this module.
 20
 21The data is located at https://doi.org/10.5281/zenodo.17384855 and released under a CC-BY-4.0 license.
 22The Zenodo record also contains a second, smaller archive ('AIR-Leish_dataset.zip'), which holds 67 of the
 23same images and is not used here.
 24
 25Please cite the publication associated with the Zenodo record if you use this dataset in your research.
 26"""
 27
 28import os
 29import json
 30import uuid
 31from tqdm import tqdm
 32from concurrent import futures
 33from typing import Union, Tuple, Optional, List, Literal
 34
 35from torch.utils.data import Dataset, DataLoader
 36
 37import torch_em
 38
 39from .. import util
 40
 41
 42URL = "https://zenodo.org/records/17384855/files/AIR_LEISH_dataset_v1.zip"
 43CHECKSUM = "36cfeeecac27cc84266c40c8676c957065f4ddca2dfe9b231e86ad84d7451d2e"
 44
 45SETS = {"set1": "Set1", "set2": "Set2"}
 46CLASS_IDS = {"amastigotes": 1, "host_cells": 2, "nuclei": 3}
 47TARGETS = ("semantic", *CLASS_IDS)
 48
 49
 50def _image_stem(file_name):
 51    stem = os.path.basename(file_name)
 52    while stem.lower().endswith(".png"):
 53        stem = stem[:-len(".png")]
 54    return stem
 55
 56
 57def _write_atomic(path, array):
 58    import imageio.v3 as imageio
 59
 60    extension = os.path.splitext(path)[1]
 61    tmp_path = f"{os.path.splitext(path)[0]}.{uuid.uuid4().hex}.incomplete{extension}"
 62    imageio.imwrite(tmp_path, array)
 63    os.replace(tmp_path, path)
 64
 65
 66def _prepare_item(image_path, rgb_path, label_path, polygons, size):
 67    import numpy as np
 68    from PIL import Image, ImageDraw
 69
 70    if not os.path.exists(rgb_path):
 71        with Image.open(image_path) as image:
 72            _write_atomic(rgb_path, np.asarray(image.convert("RGB")))
 73
 74    if label_path is None or os.path.exists(label_path):
 75        return
 76
 77    canvas = Image.new("I", size, 0)
 78    draw = ImageDraw.Draw(canvas)
 79    for instance_id, parts in enumerate(polygons, start=1):
 80        for part in parts:
 81            draw.polygon(list(zip(part[0::2], part[1::2])), fill=instance_id)
 82
 83    _write_atomic(label_path, np.asarray(canvas).astype("uint16"))
 84
 85
 86def get_air_leish_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 87    """Download the AIR-LEISH dataset.
 88
 89    Args:
 90        path: Filepath to a folder where the data is downloaded for further processing.
 91        download: Whether to download the data if it is not present.
 92
 93    Returns:
 94        Filepath where the data is downloaded.
 95    """
 96    data_dir = os.path.join(path, "AIR LEISH dataset")
 97    if os.path.exists(data_dir):
 98        return data_dir
 99
100    os.makedirs(path, exist_ok=True)
101
102    zip_path = os.path.join(path, "AIR_LEISH_dataset_v1.zip")
103    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
104    util.unzip(zip_path=zip_path, dst=path, remove=False)
105
106    assert os.path.exists(data_dir), f"The extraction of the archive did not create the expected folder in '{path}'."
107
108    return data_dir
109
110
111def get_air_leish_paths(
112    path: Union[os.PathLike, str],
113    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
114    subset: Optional[Literal["set1", "set2"]] = None,
115    download: bool = False,
116) -> Tuple[List[str], List[str]]:
117    """Get paths to the AIR-LEISH data.
118
119    Args:
120        path: Filepath to a folder where the data is downloaded for further processing.
121        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
122            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
123        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
124        download: Whether to download the data if it is not present.
125
126    Returns:
127        List of filepaths for the image data.
128        List of filepaths for the label data.
129    """
130    if target not in TARGETS:
131        raise ValueError(f"'{target}' is not a valid target. Choose one of {list(TARGETS)}.")
132    if subset is not None and subset not in SETS:
133        raise ValueError(f"'{subset}' is not a valid subset. Choose one of {list(SETS)}.")
134
135    data_dir = get_air_leish_data(path, download)
136    rgb_dir = os.path.join(path, "images_rgb")
137    label_dir = os.path.join(path, "labels", target)
138    os.makedirs(rgb_dir, exist_ok=True)
139    os.makedirs(label_dir, exist_ok=True)
140
141    jobs, raw_paths, label_paths = [], [], []
142    for set_name in [SETS[subset]] if subset is not None else SETS.values():
143        with open(os.path.join(data_dir, set_name, "_annotations.coco.json")) as f:
144            coco = json.load(f)
145
146        instances = {}
147        for annotation in coco["annotations"]:
148            if target != "semantic" and annotation["category_id"] == CLASS_IDS[target]:
149                instances.setdefault(annotation["image_id"], []).append(annotation["segmentation"])
150
151        for image in sorted(coco["images"], key=lambda entry: _image_stem(entry["file_name"])):
152            stem = _image_stem(image["file_name"])
153            image_path = os.path.join(data_dir, set_name, "Images", f"{stem}.png")
154            assert os.path.exists(image_path), f"Cannot find the image for '{image['file_name']}'."
155
156            rgb_path = os.path.join(rgb_dir, f"{set_name}_{stem}.png")
157            if target == "semantic":
158                mask_path = os.path.join(data_dir, set_name, "Masks", f"{stem}.png")
159                if not os.path.exists(mask_path):
160                    continue
161                label_path, polygons = None, None
162            else:
163                label_path = os.path.join(label_dir, f"{set_name}_{stem}.tif")
164                polygons = instances.get(image["id"], [])
165                mask_path = label_path
166
167            jobs.append((image_path, rgb_path, label_path, polygons, (image["width"], image["height"])))
168            raw_paths.append(rgb_path)
169            label_paths.append(mask_path)
170
171    with futures.ThreadPoolExecutor(min(8, os.cpu_count() or 1)) as pool:
172        tasks = [pool.submit(_prepare_item, *job) for job in jobs]
173        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Prepare AIR-LEISH"):
174            task.result()
175
176    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
177
178    return raw_paths, label_paths
179
180
181def get_air_leish_dataset(
182    path: Union[os.PathLike, str],
183    patch_shape: Tuple[int, int],
184    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
185    subset: Optional[Literal["set1", "set2"]] = None,
186    resize_inputs: bool = False,
187    download: bool = False,
188    **kwargs
189) -> Dataset:
190    """Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei.
191
192    Args:
193        path: Filepath to a folder where the data is downloaded for further processing.
194        patch_shape: The patch shape to use for training.
195        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
196            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
197        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
198        resize_inputs: Whether to resize the inputs to the patch shape.
199        download: Whether to download the data if it is not present.
200        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
201
202    Returns:
203        The segmentation dataset.
204    """
205    raw_paths, label_paths = get_air_leish_paths(path, target, subset, download)
206
207    if resize_inputs:
208        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
209        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
210            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
211        )
212
213    return torch_em.default_segmentation_dataset(
214        raw_paths=raw_paths,
215        raw_key=None,
216        label_paths=label_paths,
217        label_key=None,
218        is_seg_dataset=False,
219        patch_shape=patch_shape,
220        **kwargs
221    )
222
223
224def get_air_leish_loader(
225    path: Union[os.PathLike, str],
226    batch_size: int,
227    patch_shape: Tuple[int, int],
228    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
229    subset: Optional[Literal["set1", "set2"]] = None,
230    resize_inputs: bool = False,
231    download: bool = False,
232    **kwargs
233) -> DataLoader:
234    """Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei.
235
236    Args:
237        path: Filepath to a folder where the data is downloaded for further processing.
238        batch_size: The batch size for training.
239        patch_shape: The patch shape to use for training.
240        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
241            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
242        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
243        resize_inputs: Whether to resize the inputs to the patch shape.
244        download: Whether to download the data if it is not present.
245        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
246
247    Returns:
248        The DataLoader.
249    """
250    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
251    dataset = get_air_leish_dataset(path, patch_shape, target, subset, resize_inputs, download, **ds_kwargs)
252    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/records/17384855/files/AIR_LEISH_dataset_v1.zip'
CHECKSUM = '36cfeeecac27cc84266c40c8676c957065f4ddca2dfe9b231e86ad84d7451d2e'
SETS = {'set1': 'Set1', 'set2': 'Set2'}
CLASS_IDS = {'amastigotes': 1, 'host_cells': 2, 'nuclei': 3}
TARGETS = ('semantic', 'amastigotes', 'host_cells', 'nuclei')
def get_air_leish_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 87def get_air_leish_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 88    """Download the AIR-LEISH dataset.
 89
 90    Args:
 91        path: Filepath to a folder where the data is downloaded for further processing.
 92        download: Whether to download the data if it is not present.
 93
 94    Returns:
 95        Filepath where the data is downloaded.
 96    """
 97    data_dir = os.path.join(path, "AIR LEISH dataset")
 98    if os.path.exists(data_dir):
 99        return data_dir
100
101    os.makedirs(path, exist_ok=True)
102
103    zip_path = os.path.join(path, "AIR_LEISH_dataset_v1.zip")
104    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
105    util.unzip(zip_path=zip_path, dst=path, remove=False)
106
107    assert os.path.exists(data_dir), f"The extraction of the archive did not create the expected folder in '{path}'."
108
109    return data_dir

Download the AIR-LEISH dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_air_leish_paths( path: Union[os.PathLike, str], target: Literal['semantic', 'amastigotes', 'host_cells', 'nuclei'] = 'semantic', subset: Optional[Literal['set1', 'set2']] = None, download: bool = False) -> Tuple[List[str], List[str]]:
112def get_air_leish_paths(
113    path: Union[os.PathLike, str],
114    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
115    subset: Optional[Literal["set1", "set2"]] = None,
116    download: bool = False,
117) -> Tuple[List[str], List[str]]:
118    """Get paths to the AIR-LEISH data.
119
120    Args:
121        path: Filepath to a folder where the data is downloaded for further processing.
122        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
123            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
124        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
125        download: Whether to download the data if it is not present.
126
127    Returns:
128        List of filepaths for the image data.
129        List of filepaths for the label data.
130    """
131    if target not in TARGETS:
132        raise ValueError(f"'{target}' is not a valid target. Choose one of {list(TARGETS)}.")
133    if subset is not None and subset not in SETS:
134        raise ValueError(f"'{subset}' is not a valid subset. Choose one of {list(SETS)}.")
135
136    data_dir = get_air_leish_data(path, download)
137    rgb_dir = os.path.join(path, "images_rgb")
138    label_dir = os.path.join(path, "labels", target)
139    os.makedirs(rgb_dir, exist_ok=True)
140    os.makedirs(label_dir, exist_ok=True)
141
142    jobs, raw_paths, label_paths = [], [], []
143    for set_name in [SETS[subset]] if subset is not None else SETS.values():
144        with open(os.path.join(data_dir, set_name, "_annotations.coco.json")) as f:
145            coco = json.load(f)
146
147        instances = {}
148        for annotation in coco["annotations"]:
149            if target != "semantic" and annotation["category_id"] == CLASS_IDS[target]:
150                instances.setdefault(annotation["image_id"], []).append(annotation["segmentation"])
151
152        for image in sorted(coco["images"], key=lambda entry: _image_stem(entry["file_name"])):
153            stem = _image_stem(image["file_name"])
154            image_path = os.path.join(data_dir, set_name, "Images", f"{stem}.png")
155            assert os.path.exists(image_path), f"Cannot find the image for '{image['file_name']}'."
156
157            rgb_path = os.path.join(rgb_dir, f"{set_name}_{stem}.png")
158            if target == "semantic":
159                mask_path = os.path.join(data_dir, set_name, "Masks", f"{stem}.png")
160                if not os.path.exists(mask_path):
161                    continue
162                label_path, polygons = None, None
163            else:
164                label_path = os.path.join(label_dir, f"{set_name}_{stem}.tif")
165                polygons = instances.get(image["id"], [])
166                mask_path = label_path
167
168            jobs.append((image_path, rgb_path, label_path, polygons, (image["width"], image["height"])))
169            raw_paths.append(rgb_path)
170            label_paths.append(mask_path)
171
172    with futures.ThreadPoolExecutor(min(8, os.cpu_count() or 1)) as pool:
173        tasks = [pool.submit(_prepare_item, *job) for job in jobs]
174        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Prepare AIR-LEISH"):
175            task.result()
176
177    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
178
179    return raw_paths, label_paths

Get paths to the AIR-LEISH data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
  • subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_air_leish_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], target: Literal['semantic', 'amastigotes', 'host_cells', 'nuclei'] = 'semantic', subset: Optional[Literal['set1', 'set2']] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
182def get_air_leish_dataset(
183    path: Union[os.PathLike, str],
184    patch_shape: Tuple[int, int],
185    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
186    subset: Optional[Literal["set1", "set2"]] = None,
187    resize_inputs: bool = False,
188    download: bool = False,
189    **kwargs
190) -> Dataset:
191    """Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei.
192
193    Args:
194        path: Filepath to a folder where the data is downloaded for further processing.
195        patch_shape: The patch shape to use for training.
196        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
197            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
198        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
199        resize_inputs: Whether to resize the inputs to the patch shape.
200        download: Whether to download the data if it is not present.
201        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
202
203    Returns:
204        The segmentation dataset.
205    """
206    raw_paths, label_paths = get_air_leish_paths(path, target, subset, download)
207
208    if resize_inputs:
209        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
210        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
211            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
212        )
213
214    return torch_em.default_segmentation_dataset(
215        raw_paths=raw_paths,
216        raw_key=None,
217        label_paths=label_paths,
218        label_key=None,
219        is_seg_dataset=False,
220        patch_shape=patch_shape,
221        **kwargs
222    )

Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
  • subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_air_leish_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], target: Literal['semantic', 'amastigotes', 'host_cells', 'nuclei'] = 'semantic', subset: Optional[Literal['set1', 'set2']] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
225def get_air_leish_loader(
226    path: Union[os.PathLike, str],
227    batch_size: int,
228    patch_shape: Tuple[int, int],
229    target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic",
230    subset: Optional[Literal["set1", "set2"]] = None,
231    resize_inputs: bool = False,
232    download: bool = False,
233    **kwargs
234) -> DataLoader:
235    """Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei.
236
237    Args:
238        path: Filepath to a folder where the data is downloaded for further processing.
239        batch_size: The batch size for training.
240        patch_shape: The patch shape to use for training.
241        target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus)
242            or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
243        subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
244        resize_inputs: Whether to resize the inputs to the patch shape.
245        download: Whether to download the data if it is not present.
246        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
247
248    Returns:
249        The DataLoader.
250    """
251    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
252    dataset = get_air_leish_dataset(path, patch_shape, target, subset, resize_inputs, download, **ds_kwargs)
253    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
  • subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.