torch_em.data.datasets.light_microscopy.air_leish
The AIR-LEISH dataset contains annotations for the segmentation of Leishmania amastigotes, host cells (macrophages) and nuclei in microscopy images of Giemsa-stained Leishmania-infected macrophages.
The dataset consists of 180 RGB images of 1844 x 2709 pixels, in two sets of 90 images (Set1, file names
'20250328_CC*', and Set2, file names '20250203_CF*'). Every image is annotated with three classes,
'AM' (amastigotes), 'HC' (host cells) and 'NU' (nuclei), as COCO polygons (one polygon per instance) and,
with one exception, as a pixel-wise class mask (0: background, 1: amastigote, 2: host cell, 3: nucleus).
In the class masks the nuclei are cut out of the host cells (the polygon of a host cell covers its nucleus).
The target argument selects the labels:
- 'semantic': the class masks shipped with the dataset. Image '20250328_CCimage49' of
Set1has no class mask and is not part of this target (179 images). - 'amastigotes', 'host_cells' or 'nuclei': instance labels rasterized from the COCO polygons of this class (all 180 images, one id per instance in the order of the annotation file, later instances overwrite earlier ones where they overlap, which is rare). 26 images do not contain any amastigote and have an empty label.
NOTE: The images are stored as RGBA and are converted to RGB, and the labels of the instance targets are
rasterized once when the paths are first requested. 54 file names in the annotations of Set2 have a doubled
'.png' extension, which is handled by this module.
The data is located at https://doi.org/10.5281/zenodo.17384855 and released under a CC-BY-4.0 license. The Zenodo record also contains a second, smaller archive ('AIR-Leish_dataset.zip'), which holds 67 of the same images and is not used here.
Please cite the publication associated with the Zenodo record if you use this dataset in your research.
1"""The AIR-LEISH dataset contains annotations for the segmentation of Leishmania amastigotes, host cells 2(macrophages) and nuclei in microscopy images of Giemsa-stained Leishmania-infected macrophages. 3 4The dataset consists of 180 RGB images of 1844 x 2709 pixels, in two sets of 90 images (`Set1`, file names 5'20250328_CC*', and `Set2`, file names '20250203_CF*'). Every image is annotated with three classes, 6'AM' (amastigotes), 'HC' (host cells) and 'NU' (nuclei), as COCO polygons (one polygon per instance) and, 7with one exception, as a pixel-wise class mask (0: background, 1: amastigote, 2: host cell, 3: nucleus). 8In the class masks the nuclei are cut out of the host cells (the polygon of a host cell covers its nucleus). 9 10The `target` argument selects the labels: 11- 'semantic': the class masks shipped with the dataset. Image '20250328_CCimage49' of `Set1` has no class mask 12 and is not part of this target (179 images). 13- 'amastigotes', 'host_cells' or 'nuclei': instance labels rasterized from the COCO polygons of this class 14 (all 180 images, one id per instance in the order of the annotation file, later instances overwrite earlier 15 ones where they overlap, which is rare). 26 images do not contain any amastigote and have an empty label. 16 17NOTE: The images are stored as RGBA and are converted to RGB, and the labels of the instance targets are 18rasterized once when the paths are first requested. 54 file names in the annotations of `Set2` have a doubled 19'.png' extension, which is handled by this module. 20 21The data is located at https://doi.org/10.5281/zenodo.17384855 and released under a CC-BY-4.0 license. 22The Zenodo record also contains a second, smaller archive ('AIR-Leish_dataset.zip'), which holds 67 of the 23same images and is not used here. 24 25Please cite the publication associated with the Zenodo record if you use this dataset in your research. 26""" 27 28import os 29import json 30import uuid 31from tqdm import tqdm 32from concurrent import futures 33from typing import Union, Tuple, Optional, List, Literal 34 35from torch.utils.data import Dataset, DataLoader 36 37import torch_em 38 39from .. import util 40 41 42URL = "https://zenodo.org/records/17384855/files/AIR_LEISH_dataset_v1.zip" 43CHECKSUM = "36cfeeecac27cc84266c40c8676c957065f4ddca2dfe9b231e86ad84d7451d2e" 44 45SETS = {"set1": "Set1", "set2": "Set2"} 46CLASS_IDS = {"amastigotes": 1, "host_cells": 2, "nuclei": 3} 47TARGETS = ("semantic", *CLASS_IDS) 48 49 50def _image_stem(file_name): 51 stem = os.path.basename(file_name) 52 while stem.lower().endswith(".png"): 53 stem = stem[:-len(".png")] 54 return stem 55 56 57def _write_atomic(path, array): 58 import imageio.v3 as imageio 59 60 extension = os.path.splitext(path)[1] 61 tmp_path = f"{os.path.splitext(path)[0]}.{uuid.uuid4().hex}.incomplete{extension}" 62 imageio.imwrite(tmp_path, array) 63 os.replace(tmp_path, path) 64 65 66def _prepare_item(image_path, rgb_path, label_path, polygons, size): 67 import numpy as np 68 from PIL import Image, ImageDraw 69 70 if not os.path.exists(rgb_path): 71 with Image.open(image_path) as image: 72 _write_atomic(rgb_path, np.asarray(image.convert("RGB"))) 73 74 if label_path is None or os.path.exists(label_path): 75 return 76 77 canvas = Image.new("I", size, 0) 78 draw = ImageDraw.Draw(canvas) 79 for instance_id, parts in enumerate(polygons, start=1): 80 for part in parts: 81 draw.polygon(list(zip(part[0::2], part[1::2])), fill=instance_id) 82 83 _write_atomic(label_path, np.asarray(canvas).astype("uint16")) 84 85 86def get_air_leish_data(path: Union[os.PathLike, str], download: bool = False) -> str: 87 """Download the AIR-LEISH dataset. 88 89 Args: 90 path: Filepath to a folder where the data is downloaded for further processing. 91 download: Whether to download the data if it is not present. 92 93 Returns: 94 Filepath where the data is downloaded. 95 """ 96 data_dir = os.path.join(path, "AIR LEISH dataset") 97 if os.path.exists(data_dir): 98 return data_dir 99 100 os.makedirs(path, exist_ok=True) 101 102 zip_path = os.path.join(path, "AIR_LEISH_dataset_v1.zip") 103 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 104 util.unzip(zip_path=zip_path, dst=path, remove=False) 105 106 assert os.path.exists(data_dir), f"The extraction of the archive did not create the expected folder in '{path}'." 107 108 return data_dir 109 110 111def get_air_leish_paths( 112 path: Union[os.PathLike, str], 113 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 114 subset: Optional[Literal["set1", "set2"]] = None, 115 download: bool = False, 116) -> Tuple[List[str], List[str]]: 117 """Get paths to the AIR-LEISH data. 118 119 Args: 120 path: Filepath to a folder where the data is downloaded for further processing. 121 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 122 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 123 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 124 download: Whether to download the data if it is not present. 125 126 Returns: 127 List of filepaths for the image data. 128 List of filepaths for the label data. 129 """ 130 if target not in TARGETS: 131 raise ValueError(f"'{target}' is not a valid target. Choose one of {list(TARGETS)}.") 132 if subset is not None and subset not in SETS: 133 raise ValueError(f"'{subset}' is not a valid subset. Choose one of {list(SETS)}.") 134 135 data_dir = get_air_leish_data(path, download) 136 rgb_dir = os.path.join(path, "images_rgb") 137 label_dir = os.path.join(path, "labels", target) 138 os.makedirs(rgb_dir, exist_ok=True) 139 os.makedirs(label_dir, exist_ok=True) 140 141 jobs, raw_paths, label_paths = [], [], [] 142 for set_name in [SETS[subset]] if subset is not None else SETS.values(): 143 with open(os.path.join(data_dir, set_name, "_annotations.coco.json")) as f: 144 coco = json.load(f) 145 146 instances = {} 147 for annotation in coco["annotations"]: 148 if target != "semantic" and annotation["category_id"] == CLASS_IDS[target]: 149 instances.setdefault(annotation["image_id"], []).append(annotation["segmentation"]) 150 151 for image in sorted(coco["images"], key=lambda entry: _image_stem(entry["file_name"])): 152 stem = _image_stem(image["file_name"]) 153 image_path = os.path.join(data_dir, set_name, "Images", f"{stem}.png") 154 assert os.path.exists(image_path), f"Cannot find the image for '{image['file_name']}'." 155 156 rgb_path = os.path.join(rgb_dir, f"{set_name}_{stem}.png") 157 if target == "semantic": 158 mask_path = os.path.join(data_dir, set_name, "Masks", f"{stem}.png") 159 if not os.path.exists(mask_path): 160 continue 161 label_path, polygons = None, None 162 else: 163 label_path = os.path.join(label_dir, f"{set_name}_{stem}.tif") 164 polygons = instances.get(image["id"], []) 165 mask_path = label_path 166 167 jobs.append((image_path, rgb_path, label_path, polygons, (image["width"], image["height"]))) 168 raw_paths.append(rgb_path) 169 label_paths.append(mask_path) 170 171 with futures.ThreadPoolExecutor(min(8, os.cpu_count() or 1)) as pool: 172 tasks = [pool.submit(_prepare_item, *job) for job in jobs] 173 for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Prepare AIR-LEISH"): 174 task.result() 175 176 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 177 178 return raw_paths, label_paths 179 180 181def get_air_leish_dataset( 182 path: Union[os.PathLike, str], 183 patch_shape: Tuple[int, int], 184 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 185 subset: Optional[Literal["set1", "set2"]] = None, 186 resize_inputs: bool = False, 187 download: bool = False, 188 **kwargs 189) -> Dataset: 190 """Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei. 191 192 Args: 193 path: Filepath to a folder where the data is downloaded for further processing. 194 patch_shape: The patch shape to use for training. 195 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 196 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 197 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 198 resize_inputs: Whether to resize the inputs to the patch shape. 199 download: Whether to download the data if it is not present. 200 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 201 202 Returns: 203 The segmentation dataset. 204 """ 205 raw_paths, label_paths = get_air_leish_paths(path, target, subset, download) 206 207 if resize_inputs: 208 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 209 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 210 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 211 ) 212 213 return torch_em.default_segmentation_dataset( 214 raw_paths=raw_paths, 215 raw_key=None, 216 label_paths=label_paths, 217 label_key=None, 218 is_seg_dataset=False, 219 patch_shape=patch_shape, 220 **kwargs 221 ) 222 223 224def get_air_leish_loader( 225 path: Union[os.PathLike, str], 226 batch_size: int, 227 patch_shape: Tuple[int, int], 228 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 229 subset: Optional[Literal["set1", "set2"]] = None, 230 resize_inputs: bool = False, 231 download: bool = False, 232 **kwargs 233) -> DataLoader: 234 """Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei. 235 236 Args: 237 path: Filepath to a folder where the data is downloaded for further processing. 238 batch_size: The batch size for training. 239 patch_shape: The patch shape to use for training. 240 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 241 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 242 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 243 resize_inputs: Whether to resize the inputs to the patch shape. 244 download: Whether to download the data if it is not present. 245 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 246 247 Returns: 248 The DataLoader. 249 """ 250 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 251 dataset = get_air_leish_dataset(path, patch_shape, target, subset, resize_inputs, download, **ds_kwargs) 252 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
87def get_air_leish_data(path: Union[os.PathLike, str], download: bool = False) -> str: 88 """Download the AIR-LEISH dataset. 89 90 Args: 91 path: Filepath to a folder where the data is downloaded for further processing. 92 download: Whether to download the data if it is not present. 93 94 Returns: 95 Filepath where the data is downloaded. 96 """ 97 data_dir = os.path.join(path, "AIR LEISH dataset") 98 if os.path.exists(data_dir): 99 return data_dir 100 101 os.makedirs(path, exist_ok=True) 102 103 zip_path = os.path.join(path, "AIR_LEISH_dataset_v1.zip") 104 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 105 util.unzip(zip_path=zip_path, dst=path, remove=False) 106 107 assert os.path.exists(data_dir), f"The extraction of the archive did not create the expected folder in '{path}'." 108 109 return data_dir
Download the AIR-LEISH dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is downloaded.
112def get_air_leish_paths( 113 path: Union[os.PathLike, str], 114 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 115 subset: Optional[Literal["set1", "set2"]] = None, 116 download: bool = False, 117) -> Tuple[List[str], List[str]]: 118 """Get paths to the AIR-LEISH data. 119 120 Args: 121 path: Filepath to a folder where the data is downloaded for further processing. 122 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 123 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 124 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 125 download: Whether to download the data if it is not present. 126 127 Returns: 128 List of filepaths for the image data. 129 List of filepaths for the label data. 130 """ 131 if target not in TARGETS: 132 raise ValueError(f"'{target}' is not a valid target. Choose one of {list(TARGETS)}.") 133 if subset is not None and subset not in SETS: 134 raise ValueError(f"'{subset}' is not a valid subset. Choose one of {list(SETS)}.") 135 136 data_dir = get_air_leish_data(path, download) 137 rgb_dir = os.path.join(path, "images_rgb") 138 label_dir = os.path.join(path, "labels", target) 139 os.makedirs(rgb_dir, exist_ok=True) 140 os.makedirs(label_dir, exist_ok=True) 141 142 jobs, raw_paths, label_paths = [], [], [] 143 for set_name in [SETS[subset]] if subset is not None else SETS.values(): 144 with open(os.path.join(data_dir, set_name, "_annotations.coco.json")) as f: 145 coco = json.load(f) 146 147 instances = {} 148 for annotation in coco["annotations"]: 149 if target != "semantic" and annotation["category_id"] == CLASS_IDS[target]: 150 instances.setdefault(annotation["image_id"], []).append(annotation["segmentation"]) 151 152 for image in sorted(coco["images"], key=lambda entry: _image_stem(entry["file_name"])): 153 stem = _image_stem(image["file_name"]) 154 image_path = os.path.join(data_dir, set_name, "Images", f"{stem}.png") 155 assert os.path.exists(image_path), f"Cannot find the image for '{image['file_name']}'." 156 157 rgb_path = os.path.join(rgb_dir, f"{set_name}_{stem}.png") 158 if target == "semantic": 159 mask_path = os.path.join(data_dir, set_name, "Masks", f"{stem}.png") 160 if not os.path.exists(mask_path): 161 continue 162 label_path, polygons = None, None 163 else: 164 label_path = os.path.join(label_dir, f"{set_name}_{stem}.tif") 165 polygons = instances.get(image["id"], []) 166 mask_path = label_path 167 168 jobs.append((image_path, rgb_path, label_path, polygons, (image["width"], image["height"]))) 169 raw_paths.append(rgb_path) 170 label_paths.append(mask_path) 171 172 with futures.ThreadPoolExecutor(min(8, os.cpu_count() or 1)) as pool: 173 tasks = [pool.submit(_prepare_item, *job) for job in jobs] 174 for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Prepare AIR-LEISH"): 175 task.result() 176 177 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 178 179 return raw_paths, label_paths
Get paths to the AIR-LEISH data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
- subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
182def get_air_leish_dataset( 183 path: Union[os.PathLike, str], 184 patch_shape: Tuple[int, int], 185 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 186 subset: Optional[Literal["set1", "set2"]] = None, 187 resize_inputs: bool = False, 188 download: bool = False, 189 **kwargs 190) -> Dataset: 191 """Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei. 192 193 Args: 194 path: Filepath to a folder where the data is downloaded for further processing. 195 patch_shape: The patch shape to use for training. 196 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 197 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 198 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 199 resize_inputs: Whether to resize the inputs to the patch shape. 200 download: Whether to download the data if it is not present. 201 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 202 203 Returns: 204 The segmentation dataset. 205 """ 206 raw_paths, label_paths = get_air_leish_paths(path, target, subset, download) 207 208 if resize_inputs: 209 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 210 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 211 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 212 ) 213 214 return torch_em.default_segmentation_dataset( 215 raw_paths=raw_paths, 216 raw_key=None, 217 label_paths=label_paths, 218 label_key=None, 219 is_seg_dataset=False, 220 patch_shape=patch_shape, 221 **kwargs 222 )
Get the AIR-LEISH dataset for the segmentation of Leishmania amastigotes, host cells and nuclei.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
- subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
- resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
225def get_air_leish_loader( 226 path: Union[os.PathLike, str], 227 batch_size: int, 228 patch_shape: Tuple[int, int], 229 target: Literal["semantic", "amastigotes", "host_cells", "nuclei"] = "semantic", 230 subset: Optional[Literal["set1", "set2"]] = None, 231 resize_inputs: bool = False, 232 download: bool = False, 233 **kwargs 234) -> DataLoader: 235 """Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei. 236 237 Args: 238 path: Filepath to a folder where the data is downloaded for further processing. 239 batch_size: The batch size for training. 240 patch_shape: The patch shape to use for training. 241 target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) 242 or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'. 243 subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used. 244 resize_inputs: Whether to resize the inputs to the patch shape. 245 download: Whether to download the data if it is not present. 246 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 247 248 Returns: 249 The DataLoader. 250 """ 251 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 252 dataset = get_air_leish_dataset(path, patch_shape, target, subset, resize_inputs, download, **ds_kwargs) 253 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the AIR-LEISH dataloader for the segmentation of Leishmania amastigotes, host cells and nuclei.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- target: The choice of labels. Either 'semantic' (class masks with 1: amastigote, 2: host cell, 3: nucleus) or the instances of one class: 'amastigotes', 'host_cells' or 'nuclei'.
- subset: The choice of image set. Either 'set1' or 'set2'. By default both sets are used.
- resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.