torch_em.data.datasets.medical.hubmap_kidney

The HuBMAP - Hacking the Kidney dataset contains annotations for glomeruli functional tissue unit (FTU) segmentation in human kidney whole-slide histopathology images (PAS-stained).

This is the dataset for the "HuBMAP - Hacking the Kidney" Kaggle competition, located at https://www.kaggle.com/competitions/hubmap-kidney-segmentation. It provides 8 PAS-stained kidney whole-slide images (fresh-frozen and FFPE tissue preparations, contributed by the Human BioMolecular Atlas Program / HuBMAP through the BIOMIC team at Vanderbilt University), each paired with expert-annotated glomerulus segmentations, given as run-length encoded (RLE) masks in 'train.csv'. The competition data is released under the CC BY 4.0 license.

NOTE: This is a DIFFERENT Kaggle competition from "HuBMAP + HPA - Hacking the Human Body" (see 'hubmap_hpa.py' for that dataset), even though both are hosted by HuBMAP and target FTU segmentation.

NOTE: Downloading this dataset requires a Kaggle account that has accepted the competition rules at https://www.kaggle.com/competitions/hubmap-kidney-segmentation/rules. Without that, the Kaggle API download fails with an HTTP 403 error, even with valid API credentials.

NOTE: The whole-slide images are very large single-resolution BigTIFFs (tens of thousands of pixels per side, several hundred MB to multiple GB per file). As in 'histopathology/camelyon.py', each image is read lazily (via a 'zarr' view over the tiled TIFF where available, otherwise a memory-mapped read) and converted once into a chunked HDF5 file that stores the raw image together with the binary glomerulus mask decoded from the RLE encoding. Downstream patch extraction for training then happens lazily via 'torch_em.default_segmentation_dataset' from these HDF5 files, so no manual tiling logic is required here.

This dataset is described in the publication https://doi.org/10.1038/s41467-023-40291-0 ("Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms", Nature Communications, 2023). Please cite it if you use this dataset in your research.

  1"""The HuBMAP - Hacking the Kidney dataset contains annotations for glomeruli functional tissue unit (FTU)
  2segmentation in human kidney whole-slide histopathology images (PAS-stained).
  3
  4This is the dataset for the "HuBMAP - Hacking the Kidney" Kaggle competition, located at
  5https://www.kaggle.com/competitions/hubmap-kidney-segmentation. It provides 8 PAS-stained kidney whole-slide
  6images (fresh-frozen and FFPE tissue preparations, contributed by the Human BioMolecular Atlas Program /
  7HuBMAP through the BIOMIC team at Vanderbilt University), each paired with expert-annotated glomerulus
  8segmentations, given as run-length encoded (RLE) masks in 'train.csv'. The competition data is released
  9under the CC BY 4.0 license.
 10
 11NOTE: This is a DIFFERENT Kaggle competition from "HuBMAP + HPA - Hacking the Human Body" (see
 12'hubmap_hpa.py' for that dataset), even though both are hosted by HuBMAP and target FTU segmentation.
 13
 14NOTE: Downloading this dataset requires a Kaggle account that has accepted the competition rules at
 15https://www.kaggle.com/competitions/hubmap-kidney-segmentation/rules. Without that, the Kaggle API
 16download fails with an HTTP 403 error, even with valid API credentials.
 17
 18NOTE: The whole-slide images are very large single-resolution BigTIFFs (tens of thousands of pixels per
 19side, several hundred MB to multiple GB per file). As in 'histopathology/camelyon.py', each image is read
 20lazily (via a 'zarr' view over the tiled TIFF where available, otherwise a memory-mapped read) and converted
 21once into a chunked HDF5 file that stores the raw image together with the binary glomerulus mask decoded
 22from the RLE encoding. Downstream patch extraction for training then happens lazily via
 23'torch_em.default_segmentation_dataset' from these HDF5 files, so no manual tiling logic is required here.
 24
 25This dataset is described in the publication https://doi.org/10.1038/s41467-023-40291-0 ("Segmenting
 26functional tissue units across human organs using community-driven development of generalizable machine
 27learning algorithms", Nature Communications, 2023). Please cite it if you use this dataset in your research.
 28"""
 29
 30import os
 31import csv
 32from pathlib import Path
 33from typing import List, Tuple, Union
 34
 35import numpy as np
 36from tqdm import tqdm
 37
 38from torch.utils.data import Dataset, DataLoader
 39
 40import torch_em
 41
 42from .. import util
 43
 44# The RLE-encoded masks in 'train.csv' exceed the default field size limit for these large WSIs.
 45csv.field_size_limit(10**8)
 46
 47
 48def _rle_decode(rle: str, shape: Tuple[int, int]) -> np.ndarray:
 49    """Decode a Kaggle-style run-length-encoded mask into a binary array.
 50
 51    The encoding lists 1-indexed (start, length) pairs over a column-major (Fortran order) flattening of
 52    the image, i.e. pixels are numbered from top to bottom, then left to right.
 53    """
 54    height, width = shape
 55    mask = np.zeros(height * width, dtype=np.uint8)
 56
 57    values = [int(v) for v in rle.split()]
 58    starts, lengths = values[0::2], values[1::2]
 59    for start, length in zip(starts, lengths):
 60        start -= 1
 61        mask[start:start + length] = 1
 62
 63    return mask.reshape((width, height)).T
 64
 65
 66def _open_raw_array(tiff_path):
 67    import tifffile
 68
 69    tiff = tifffile.TiffFile(tiff_path)
 70    # Some slides are stored as several series (e.g. a full-resolution image plus a thumbnail);
 71    # pick the one with the most pixels rather than assuming series[0] is the full-resolution one.
 72    series = max(tiff.series, key=lambda s: np.prod(s.shape))
 73
 74    try:
 75        import zarr
 76        array = zarr.open(series.aszarr(), mode="r")
 77        if not hasattr(array, "shape"):  # The tiff is pyramidal: pick the largest resolution level.
 78            array = max(array.values(), key=lambda level: np.prod(level.shape))
 79    except Exception:
 80        array = series.asarray()
 81
 82    # Some slides carry extra singleton axes (e.g. an OME-style (T, Z, C, H, W) layout), which have
 83    # to be dropped before the channel axis can be identified. Indexed rather than via '.squeeze()',
 84    # so a lazy zarr array is not fully materialized just to drop size-1 axes.
 85    squeeze_axes = tuple(i for i, size in enumerate(array.shape) if size == 1)
 86    if squeeze_axes:
 87        index = tuple(0 if i in squeeze_axes else slice(None) for i in range(array.ndim))
 88        array = array[index]
 89
 90    # Some HuBMAP kidney tiffs store the channel axis first, i.e. (3, height, width) instead of the usual
 91    # channel-last (height, width, 3) layout. Normalize to channel-last so downstream code is uniform.
 92    if array.ndim == 3 and array.shape[0] == 3 and array.shape[-1] != 3:
 93        array = np.transpose(array, (1, 2, 0))
 94
 95    assert array.ndim == 3 and array.shape[-1] == 3, f"Unexpected array shape {array.shape} for {tiff_path}."
 96    return array
 97
 98
 99def _convert_slide(tiff_path: str, rle: str, output_path: str, tile: int = 4096) -> None:
100    import h5py
101
102    array = _open_raw_array(tiff_path)
103    height, width = array.shape[0], array.shape[1]
104    mask = _rle_decode(rle, (height, width))
105
106    tmp_path = output_path + ".tmp"
107    with h5py.File(tmp_path, "w") as f:
108        raw = f.create_dataset(
109            "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(1, 512, 512)
110        )
111        labels = f.create_dataset(
112            "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512)
113        )
114        for y in tqdm(range(0, height, tile), desc=f"Converting {Path(tiff_path).stem}"):
115            for x in range(0, width, tile):
116                th, tw = min(tile, height - y), min(tile, width - x)
117                tile_data = np.asarray(array[y:y + th, x:x + tw])
118                raw[:, y:y + th, x:x + tw] = tile_data.transpose(2, 0, 1)
119                labels[y:y + th, x:x + tw] = mask[y:y + th, x:x + tw]
120
121    os.replace(tmp_path, output_path)
122
123
124def _load_train_annotations(data_dir):
125    csv_path = os.path.join(data_dir, "train.csv")
126    if not os.path.exists(csv_path):
127        raise RuntimeError(f"Could not find 'train.csv' at {csv_path}.")
128    with open(csv_path, "r") as f:
129        return list(csv.DictReader(f))
130
131
132def get_hubmap_kidney_data(path: Union[os.PathLike, str], download: bool = False) -> str:
133    """Download the HuBMAP - Hacking the Kidney dataset.
134
135    Args:
136        path: Filepath to a folder where the data will be saved.
137        download: Whether to download the data if it is not present.
138
139    Returns:
140        Filepath to the folder where the raw data is stored.
141    """
142    data_dir = os.path.join(path, "train")
143    if os.path.exists(data_dir):
144        return path
145
146    os.makedirs(path, exist_ok=True)
147    zip_path = os.path.join(path, "hubmap-kidney-segmentation.zip")
148    util.download_source_kaggle(
149        path=path, dataset_name="hubmap-kidney-segmentation", download=download, competition=True
150    )
151    util.unzip(zip_path=zip_path, dst=path)
152
153    if not os.path.exists(data_dir):
154        raise RuntimeError(f"The expected 'train' folder is missing after extraction at {path}.")
155    return path
156
157
158def get_hubmap_kidney_paths(path: Union[os.PathLike, str], download: bool = False) -> List[str]:
159    """Get paths to the preprocessed HuBMAP - Hacking the Kidney data.
160
161    Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the
162    binary glomerulus segmentation mask under 'labels' (height, width).
163
164    Args:
165        path: Filepath to a folder where the data will be saved.
166        download: Whether to download the data if it is not present.
167
168    Returns:
169        List of filepaths to the preprocessed HDF5 files.
170    """
171    data_dir = get_hubmap_kidney_data(path, download)
172    rows = _load_train_annotations(data_dir)
173
174    preprocessed_dir = os.path.join(path, "preprocessed")
175    os.makedirs(preprocessed_dir, exist_ok=True)
176
177    volume_paths = []
178    for row in rows:
179        image_id = row["id"]
180        tiff_path = os.path.join(data_dir, "train", f"{image_id}.tiff")
181        if not os.path.exists(tiff_path):
182            continue
183
184        output_path = os.path.join(preprocessed_dir, f"{image_id}.h5")
185        if not os.path.exists(output_path):
186            _convert_slide(tiff_path, row["encoding"], output_path)
187        volume_paths.append(output_path)
188
189    if not volume_paths:
190        raise RuntimeError(f"No annotated images were found at {data_dir}.")
191
192    return volume_paths
193
194
195def get_hubmap_kidney_dataset(
196    path: Union[os.PathLike, str],
197    patch_shape: Tuple[int, int],
198    download: bool = False,
199    resize_inputs: bool = False,
200    **kwargs,
201) -> Dataset:
202    """Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images.
203
204    Args:
205        path: Filepath to a folder where the data will be saved.
206        patch_shape: The patch shape to use for training.
207        download: Whether to download the data if it is not present.
208        resize_inputs: Whether to resize the inputs.
209        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
210
211    Returns:
212        The segmentation dataset.
213    """
214    volume_paths = get_hubmap_kidney_paths(path, download)
215
216    if resize_inputs:
217        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
218        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
219            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
220        )
221
222    return torch_em.default_segmentation_dataset(
223        raw_paths=volume_paths,
224        raw_key="raw",
225        label_paths=volume_paths,
226        label_key="labels",
227        patch_shape=patch_shape,
228        is_seg_dataset=True,
229        with_channels=True,
230        ndim=2,
231        **kwargs,
232    )
233
234
235def get_hubmap_kidney_loader(
236    path: Union[os.PathLike, str],
237    batch_size: int,
238    patch_shape: Tuple[int, int],
239    download: bool = False,
240    resize_inputs: bool = False,
241    **kwargs,
242) -> DataLoader:
243    """Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images.
244
245    Args:
246        path: Filepath to a folder where the data will be saved.
247        batch_size: The batch size for training.
248        patch_shape: The patch shape to use for training.
249        download: Whether to download the data if it is not present.
250        resize_inputs: Whether to resize the inputs.
251        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader.
252
253    Returns:
254        The DataLoader.
255    """
256    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
257    dataset = get_hubmap_kidney_dataset(path, patch_shape, download, resize_inputs, **ds_kwargs)
258    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
def get_hubmap_kidney_data(path: Union[os.PathLike, str], download: bool = False) -> str:
133def get_hubmap_kidney_data(path: Union[os.PathLike, str], download: bool = False) -> str:
134    """Download the HuBMAP - Hacking the Kidney dataset.
135
136    Args:
137        path: Filepath to a folder where the data will be saved.
138        download: Whether to download the data if it is not present.
139
140    Returns:
141        Filepath to the folder where the raw data is stored.
142    """
143    data_dir = os.path.join(path, "train")
144    if os.path.exists(data_dir):
145        return path
146
147    os.makedirs(path, exist_ok=True)
148    zip_path = os.path.join(path, "hubmap-kidney-segmentation.zip")
149    util.download_source_kaggle(
150        path=path, dataset_name="hubmap-kidney-segmentation", download=download, competition=True
151    )
152    util.unzip(zip_path=zip_path, dst=path)
153
154    if not os.path.exists(data_dir):
155        raise RuntimeError(f"The expected 'train' folder is missing after extraction at {path}.")
156    return path

Download the HuBMAP - Hacking the Kidney dataset.

Arguments:
  • path: Filepath to a folder where the data will be saved.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the raw data is stored.

def get_hubmap_kidney_paths(path: Union[os.PathLike, str], download: bool = False) -> List[str]:
159def get_hubmap_kidney_paths(path: Union[os.PathLike, str], download: bool = False) -> List[str]:
160    """Get paths to the preprocessed HuBMAP - Hacking the Kidney data.
161
162    Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the
163    binary glomerulus segmentation mask under 'labels' (height, width).
164
165    Args:
166        path: Filepath to a folder where the data will be saved.
167        download: Whether to download the data if it is not present.
168
169    Returns:
170        List of filepaths to the preprocessed HDF5 files.
171    """
172    data_dir = get_hubmap_kidney_data(path, download)
173    rows = _load_train_annotations(data_dir)
174
175    preprocessed_dir = os.path.join(path, "preprocessed")
176    os.makedirs(preprocessed_dir, exist_ok=True)
177
178    volume_paths = []
179    for row in rows:
180        image_id = row["id"]
181        tiff_path = os.path.join(data_dir, "train", f"{image_id}.tiff")
182        if not os.path.exists(tiff_path):
183            continue
184
185        output_path = os.path.join(preprocessed_dir, f"{image_id}.h5")
186        if not os.path.exists(output_path):
187            _convert_slide(tiff_path, row["encoding"], output_path)
188        volume_paths.append(output_path)
189
190    if not volume_paths:
191        raise RuntimeError(f"No annotated images were found at {data_dir}.")
192
193    return volume_paths

Get paths to the preprocessed HuBMAP - Hacking the Kidney data.

Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the binary glomerulus segmentation mask under 'labels' (height, width).

Arguments:
  • path: Filepath to a folder where the data will be saved.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths to the preprocessed HDF5 files.

def get_hubmap_kidney_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], download: bool = False, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
196def get_hubmap_kidney_dataset(
197    path: Union[os.PathLike, str],
198    patch_shape: Tuple[int, int],
199    download: bool = False,
200    resize_inputs: bool = False,
201    **kwargs,
202) -> Dataset:
203    """Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images.
204
205    Args:
206        path: Filepath to a folder where the data will be saved.
207        patch_shape: The patch shape to use for training.
208        download: Whether to download the data if it is not present.
209        resize_inputs: Whether to resize the inputs.
210        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
211
212    Returns:
213        The segmentation dataset.
214    """
215    volume_paths = get_hubmap_kidney_paths(path, download)
216
217    if resize_inputs:
218        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
219        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
220            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
221        )
222
223    return torch_em.default_segmentation_dataset(
224        raw_paths=volume_paths,
225        raw_key="raw",
226        label_paths=volume_paths,
227        label_key="labels",
228        patch_shape=patch_shape,
229        is_seg_dataset=True,
230        with_channels=True,
231        ndim=2,
232        **kwargs,
233    )

Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images.

Arguments:
  • path: Filepath to a folder where the data will be saved.
  • patch_shape: The patch shape to use for training.
  • download: Whether to download the data if it is not present.
  • resize_inputs: Whether to resize the inputs.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_hubmap_kidney_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], download: bool = False, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
236def get_hubmap_kidney_loader(
237    path: Union[os.PathLike, str],
238    batch_size: int,
239    patch_shape: Tuple[int, int],
240    download: bool = False,
241    resize_inputs: bool = False,
242    **kwargs,
243) -> DataLoader:
244    """Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images.
245
246    Args:
247        path: Filepath to a folder where the data will be saved.
248        batch_size: The batch size for training.
249        patch_shape: The patch shape to use for training.
250        download: Whether to download the data if it is not present.
251        resize_inputs: Whether to resize the inputs.
252        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader.
253
254    Returns:
255        The DataLoader.
256    """
257    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
258    dataset = get_hubmap_kidney_dataset(path, patch_shape, download, resize_inputs, **ds_kwargs)
259    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images.

Arguments:
  • path: Filepath to a folder where the data will be saved.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • download: Whether to download the data if it is not present.
  • resize_inputs: Whether to resize the inputs.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or the PyTorch DataLoader.
Returns:

The DataLoader.