torch_em.data.datasets.histopathology.prostate_cnb

The Prostate CNB dataset contains annotations for the semantic segmentation of prostate tissue types and histopathological structures in H&E whole-slide images of image-guided prostate core needle biopsies (CNBs).

The dataset consists of 37 whole-slide images (WSIs) with 60 core needle biopsies from 32 patients of a single institution, scanned at 0.243 micrometer per pixel. There are 15,480 manual annotations of two uropathologists (with a consensus review), which are provided as multiclass semantic masks (and as GeoJSON vectors, which are not used by this module). Regions that were not annotated are labeled as stroma automatically.

The label ids of the masks are:

  • 0: background, 1: tumor, 2: benign gland, 3: blood vessels, 4: fibromuscular bundles, 5: abnormal secretions, 6: contamination with another tissue, 7: prominent nucleolus, 8: immune cells, 9: nerve, 10: artifact, 11: seminal vesicle, 12: adipose tissue, 13: normal secretions, 14: stromal retraction spaces, 15: muscle, 16: foreign body contamination, 17: high-grade prostatic intraepithelial neoplasia (HGPIN), 18: calcifications, 19: intestinal glands and mucus, 20: perineural invasion, 21: hemorrhage, 22: intraductal carcinoma, 23: necrosis, 24: mitosis, 25: nerve ganglion, 26: atypical intraductal proliferation, 27: red blood cells, 28: stroma (see CLASS_NAMES). Not every class occurs in every slide.

NOTE: This is a different dataset than the already integrated torch_em.data.datasets.histopathology.precise (PRECISE), which consists of paired H&E and IHC slides of 25 other patients with 7 classes.

The data is located at https://doi.org/10.5281/zenodo.18930299 (the latest version is https://zenodo.org/records/21780258, the first version https://zenodo.org/records/18930300) as a single 34.5 GB zip archive with pyramidal OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole archive, this module reads its directory from the server and fetches only the requested slides and masks. Each slide is converted once into a chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the default level 1 is 0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use n_cases to only prepare a subset.

This dataset is from the publication https://doi.org/10.1038/s41597-026-08349-y. Please cite it if you use this dataset in your research.

  1"""The Prostate CNB dataset contains annotations for the semantic segmentation of prostate tissue types and
  2histopathological structures in H&E whole-slide images of image-guided prostate core needle biopsies (CNBs).
  3
  4The dataset consists of 37 whole-slide images (WSIs) with 60 core needle biopsies from 32 patients of a single
  5institution, scanned at 0.243 micrometer per pixel. There are 15,480 manual annotations of two uropathologists
  6(with a consensus review), which are provided as multiclass semantic masks (and as GeoJSON vectors, which are not used
  7by this module). Regions that were not annotated are labeled as stroma automatically.
  8
  9The label ids of the masks are:
 10- 0: background, 1: tumor, 2: benign gland, 3: blood vessels, 4: fibromuscular bundles, 5: abnormal secretions,
 11  6: contamination with another tissue, 7: prominent nucleolus, 8: immune cells, 9: nerve, 10: artifact,
 12  11: seminal vesicle, 12: adipose tissue, 13: normal secretions, 14: stromal retraction spaces, 15: muscle,
 13  16: foreign body contamination, 17: high-grade prostatic intraepithelial neoplasia (HGPIN), 18: calcifications,
 14  19: intestinal glands and mucus, 20: perineural invasion, 21: hemorrhage, 22: intraductal carcinoma,
 15  23: necrosis, 24: mitosis, 25: nerve ganglion, 26: atypical intraductal proliferation, 27: red blood cells,
 16  28: stroma
 17(see `CLASS_NAMES`). Not every class occurs in every slide.
 18
 19NOTE: This is a different dataset than the already integrated `torch_em.data.datasets.histopathology.precise`
 20(PRECISE), which consists of paired H&E and IHC slides of 25 other patients with 7 classes.
 21
 22The data is located at https://doi.org/10.5281/zenodo.18930299 (the latest version is
 23https://zenodo.org/records/21780258, the first version https://zenodo.org/records/18930300) as a single 34.5 GB zip
 24archive with pyramidal OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole
 25archive, this module reads its directory from the server and fetches only the requested slides and masks. Each slide
 26is converted once into a chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the
 27default level 1 is 0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use `n_cases` to only prepare
 28a subset.
 29
 30This dataset is from the publication https://doi.org/10.1038/s41597-026-08349-y.
 31Please cite it if you use this dataset in your research.
 32"""
 33
 34import os
 35import re
 36import json
 37import uuid
 38import zlib
 39import struct
 40from typing import List, Optional, Tuple, Union
 41
 42from tqdm import tqdm
 43
 44import torch
 45
 46from torch.utils.data import Dataset, DataLoader
 47
 48import torch_em
 49
 50from .. import util
 51
 52
 53ZIP_URL = "https://zenodo.org/api/records/21780258/files/dataset.zip/content"
 54
 55CLASS_NAMES = {
 56    0: "background",
 57    1: "tumor",
 58    2: "benign gland",
 59    3: "blood vessels",
 60    4: "fibromuscular bundles",
 61    5: "abnormal secretions",
 62    6: "contamination with another tissue",
 63    7: "prominent nucleolus",
 64    8: "immune cells",
 65    9: "nerve",
 66    10: "artifact",
 67    11: "seminal vesicle",
 68    12: "adipose tissue",
 69    13: "normal secretions",
 70    14: "stromal retraction spaces",
 71    15: "muscle",
 72    16: "foreign body contamination",
 73    17: "high-grade prostatic intraepithelial neoplasia",
 74    18: "calcifications",
 75    19: "intestinal glands and mucus",
 76    20: "perineural invasion",
 77    21: "hemorrhage",
 78    22: "intraductal carcinoma",
 79    23: "necrosis",
 80    24: "mitosis",
 81    25: "nerve ganglion",
 82    26: "atypical intraductal proliferation",
 83    27: "red blood cells",
 84    28: "stroma",
 85}
 86
 87N_LEVELS = 6
 88
 89
 90def _range_get(url, start, end, **kwargs):
 91    import requests
 92
 93    response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, **kwargs)
 94    response.raise_for_status()
 95    return response
 96
 97
 98def _read_zip_entries(path):
 99    cache_path = os.path.join(path, "zip_entries.json")
100    if os.path.exists(cache_path):
101        with open(cache_path) as f:
102            return json.load(f)
103
104    import requests
105
106    size = int(requests.head(ZIP_URL, allow_redirects=True).headers["Content-Length"])
107    tail = _range_get(ZIP_URL, size - 65557, size - 1).content
108    eocd = tail[tail.rfind(b"PK\x05\x06"):]
109    n_entries, cd_size, cd_offset = struct.unpack("<HII", eocd[10:20])
110    if cd_offset == 0xFFFFFFFF or n_entries == 0xFFFF:
111        locator = tail[tail.rfind(b"PK\x06\x07"):][:20]
112        zip64_offset = struct.unpack("<Q", locator[8:16])[0]
113        record = _range_get(ZIP_URL, zip64_offset, zip64_offset + 55).content
114        n_entries, cd_size, cd_offset = struct.unpack("<QQQ", record[32:56])
115
116    directory = _range_get(ZIP_URL, cd_offset, cd_offset + cd_size - 1).content
117    entries, pos = {}, 0
118    while pos < len(directory):
119        (signature, _, _, _, _, _, _, crc, comp_size, size, name_len, extra_len, comment_len, _, _, _,
120         header_offset) = struct.unpack("<IHHHHHHIIIHHHHHII", directory[pos:pos + 46])
121        assert signature == 0x02014B50, "The zip directory could not be parsed."
122        name = directory[pos + 46:pos + 46 + name_len].decode()
123        extra = directory[pos + 46 + name_len:pos + 46 + name_len + extra_len]
124        extra_pos = 0
125        while extra_pos < len(extra):
126            field_id, field_size = struct.unpack("<HH", extra[extra_pos:extra_pos + 4])
127            if field_id == 1:
128                field, k = extra[extra_pos + 4:extra_pos + 4 + field_size], 0
129                if size == 0xFFFFFFFF:
130                    size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8
131                if comp_size == 0xFFFFFFFF:
132                    comp_size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8
133                if header_offset == 0xFFFFFFFF:
134                    header_offset = struct.unpack("<Q", field[k:k + 8])[0]
135            extra_pos += 4 + field_size
136        if size > 0:
137            entries[name] = {"crc": crc, "comp_size": comp_size, "size": size, "offset": header_offset}
138        pos += 46 + name_len + extra_len + comment_len
139
140    tmp_path = f"{cache_path}.{uuid.uuid4().hex}.tmp"
141    with open(tmp_path, "w") as f:
142        json.dump(entries, f)
143    os.replace(tmp_path, cache_path)
144    return entries
145
146
147def _fetch_member(entry, out_path):
148    header = _range_get(ZIP_URL, entry["offset"], entry["offset"] + 29).content
149    name_len, extra_len = struct.unpack("<HH", header[26:30])
150    start = entry["offset"] + 30 + name_len + extra_len
151
152    decompressor, crc, n_bytes = zlib.decompressobj(-15), 0, 0
153    tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp"
154    with _range_get(ZIP_URL, start, start + entry["comp_size"] - 1, stream=True) as response, \
155            open(tmp_path, "wb") as f:
156        for chunk in response.iter_content(1 << 20):
157            data = decompressor.decompress(chunk)
158            crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data)
159            f.write(data)
160        data = decompressor.flush()
161        crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data)
162        f.write(data)
163
164    if crc != entry["crc"] or n_bytes != entry["size"]:
165        os.remove(tmp_path)
166        raise RuntimeError(f"The download of {os.path.basename(out_path)} is corrupted, please try again.")
167    os.replace(tmp_path, out_path)
168
169
170def _get_cases(entries):
171    pattern = re.compile(r"^images/(.+)\.ome\.tif$")
172    cases = {}
173    for name in entries:
174        match = pattern.match(name)
175        if match:
176            case = match.group(1)
177            mask_name = f"semantic_masks/{case}__mask_multiclass.ome.tif"
178            assert mask_name in entries, f"Cannot find the mask for {name}."
179            cases[case] = (name, mask_name)
180    return dict(sorted(cases.items()))
181
182
183def _open_level(series, level_index):
184    import zarr
185
186    # The pyramidal TIFFs are tiled, so the zarr view reads only the requested tiles.
187    array = zarr.open(series.aszarr(), mode="r")
188    return array if hasattr(array, "shape") else array[str(level_index)]
189
190
191def _convert_case(image_path, mask_path, out_path, level, tile=4096):
192    import h5py
193    import tifffile
194    import numpy as np
195
196    image_series = tifffile.TiffFile(image_path).series[0]
197    mask_series = tifffile.TiffFile(mask_path).series[0]
198    image, mask = _open_level(image_series, level), _open_level(mask_series, level)
199
200    # The pyramid levels of the image and mask are rounded differently, so the mask level can be a few pixels off.
201    # It is mapped to the image grid by nearest neighbor sampling, unless the shapes really differ.
202    height, width = image.shape[:2]
203    if any(abs(image.shape[i] - mask.shape[i]) > 1e-3 * image.shape[i] + 1 for i in range(2)):
204        raise RuntimeError(f"The image {image.shape} and mask {mask.shape} shapes of {image_path} do not match.")
205    rows = np.minimum(((np.arange(height) + 0.5) * mask.shape[0] / height).astype(int), mask.shape[0] - 1)
206    cols = np.minimum(((np.arange(width) + 0.5) * mask.shape[1] / width).astype(int), mask.shape[1] - 1)
207
208    tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp"
209    with h5py.File(tmp_path, "w") as f:
210        raw = f.create_dataset(
211            "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(3, 512, 512)
212        )
213        labels = f.create_dataset(
214            "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512)
215        )
216        for y in range(0, height, tile):
217            for x in range(0, width, tile):
218                bb = (slice(y, min(y + tile, height)), slice(x, min(x + tile, width)))
219                raw[(slice(None),) + bb] = image[bb].transpose(2, 0, 1)
220                r, c = rows[bb[0]], cols[bb[1]]
221                mask_block = mask[r[0]:r[-1] + 1, c[0]:c[-1] + 1]
222                labels[bb] = mask_block[np.ix_(r - r[0], c - c[0])]
223    os.replace(tmp_path, out_path)
224
225
226def get_prostate_cnb_data(
227    path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False,
228) -> str:
229    """Download and preprocess the Prostate CNB dataset.
230
231    NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched.
232    Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download.
233
234    Args:
235        path: Filepath to a folder where the data is downloaded for further processing.
236        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
237            By default all 37 slides are used.
238        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
239            micrometer. The preprocessed data is stored separately for each level.
240        download: Whether to download the data if it is not present.
241
242    Returns:
243        Filepath to the folder where the preprocessed data is stored.
244    """
245    if level not in range(N_LEVELS):
246        raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.")
247
248    os.makedirs(path, exist_ok=True)
249    if not download and not os.path.exists(os.path.join(path, "zip_entries.json")):
250        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
251    entries = _read_zip_entries(path)
252    cases = _get_cases(entries)
253    case_ids = list(cases)[:n_cases]
254
255    preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}")
256    raw_dir = os.path.join(path, "raw")
257    os.makedirs(preprocessed_dir, exist_ok=True)
258    os.makedirs(raw_dir, exist_ok=True)
259
260    missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))]
261    if missing and not download:
262        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
263
264    for case in tqdm(missing, desc="Prepare Prostate CNB slides"):
265        image_name, mask_name = cases[case]
266        image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name))
267        for name, out_path in ((image_name, image_path), (mask_name, mask_path)):
268            if not os.path.exists(out_path):
269                _fetch_member(entries[name], out_path)
270
271        _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level)
272        os.remove(image_path)
273        os.remove(mask_path)
274
275    return preprocessed_dir
276
277
278def get_prostate_cnb_paths(
279    path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False,
280) -> List[str]:
281    """Get paths to the Prostate CNB data.
282
283    Args:
284        path: Filepath to a folder where the data is downloaded for further processing.
285        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
286            By default all 37 slides are used.
287        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
288            micrometer.
289        download: Whether to download the data if it is not present.
290
291    Returns:
292        List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the
293        label data ('labels').
294    """
295    preprocessed_dir = get_prostate_cnb_data(path, n_cases, level, download)
296    case_ids = list(_get_cases(_read_zip_entries(path)))[:n_cases]
297    return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]
298
299
300def get_prostate_cnb_dataset(
301    path: Union[os.PathLike, str],
302    patch_shape: Tuple[int, int],
303    n_cases: Optional[int] = None,
304    level: int = 1,
305    download: bool = False,
306    label_dtype: torch.dtype = torch.int64,
307    resize_inputs: bool = False,
308    **kwargs
309) -> Dataset:
310    """Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images.
311
312    Args:
313        path: Filepath to a folder where the data is downloaded for further processing.
314        patch_shape: The patch shape to use for training.
315        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
316            By default all 37 slides are used.
317        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
318            micrometer.
319        download: Whether to download the data if it is not present.
320        label_dtype: The datatype of the labels.
321        resize_inputs: Whether to resize the input images.
322        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
323
324    Returns:
325        The segmentation dataset.
326    """
327    volume_paths = get_prostate_cnb_paths(path, n_cases, level, download)
328
329    if resize_inputs:
330        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
331        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
332            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
333        )
334
335    return torch_em.default_segmentation_dataset(
336        raw_paths=volume_paths,
337        raw_key="raw",
338        label_paths=volume_paths,
339        label_key="labels",
340        patch_shape=patch_shape,
341        label_dtype=label_dtype,
342        is_seg_dataset=True,
343        with_channels=True,
344        ndim=2,
345        **kwargs
346    )
347
348
349def get_prostate_cnb_loader(
350    path: Union[os.PathLike, str],
351    batch_size: int,
352    patch_shape: Tuple[int, int],
353    n_cases: Optional[int] = None,
354    level: int = 1,
355    download: bool = False,
356    label_dtype: torch.dtype = torch.int64,
357    resize_inputs: bool = False,
358    **kwargs
359) -> DataLoader:
360    """Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images.
361
362    Args:
363        path: Filepath to a folder where the data is downloaded for further processing.
364        batch_size: The batch size for training.
365        patch_shape: The patch shape to use for training.
366        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
367            By default all 37 slides are used.
368        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
369            micrometer.
370        download: Whether to download the data if it is not present.
371        label_dtype: The datatype of the labels.
372        resize_inputs: Whether to resize the input images.
373        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
374
375    Returns:
376        The DataLoader.
377    """
378    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
379    dataset = get_prostate_cnb_dataset(
380        path=path, patch_shape=patch_shape, n_cases=n_cases, level=level, download=download,
381        label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs
382    )
383    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
ZIP_URL = 'https://zenodo.org/api/records/21780258/files/dataset.zip/content'
CLASS_NAMES = {0: 'background', 1: 'tumor', 2: 'benign gland', 3: 'blood vessels', 4: 'fibromuscular bundles', 5: 'abnormal secretions', 6: 'contamination with another tissue', 7: 'prominent nucleolus', 8: 'immune cells', 9: 'nerve', 10: 'artifact', 11: 'seminal vesicle', 12: 'adipose tissue', 13: 'normal secretions', 14: 'stromal retraction spaces', 15: 'muscle', 16: 'foreign body contamination', 17: 'high-grade prostatic intraepithelial neoplasia', 18: 'calcifications', 19: 'intestinal glands and mucus', 20: 'perineural invasion', 21: 'hemorrhage', 22: 'intraductal carcinoma', 23: 'necrosis', 24: 'mitosis', 25: 'nerve ganglion', 26: 'atypical intraductal proliferation', 27: 'red blood cells', 28: 'stroma'}
N_LEVELS = 6
def get_prostate_cnb_data( path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False) -> str:
227def get_prostate_cnb_data(
228    path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False,
229) -> str:
230    """Download and preprocess the Prostate CNB dataset.
231
232    NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched.
233    Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download.
234
235    Args:
236        path: Filepath to a folder where the data is downloaded for further processing.
237        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
238            By default all 37 slides are used.
239        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
240            micrometer. The preprocessed data is stored separately for each level.
241        download: Whether to download the data if it is not present.
242
243    Returns:
244        Filepath to the folder where the preprocessed data is stored.
245    """
246    if level not in range(N_LEVELS):
247        raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.")
248
249    os.makedirs(path, exist_ok=True)
250    if not download and not os.path.exists(os.path.join(path, "zip_entries.json")):
251        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
252    entries = _read_zip_entries(path)
253    cases = _get_cases(entries)
254    case_ids = list(cases)[:n_cases]
255
256    preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}")
257    raw_dir = os.path.join(path, "raw")
258    os.makedirs(preprocessed_dir, exist_ok=True)
259    os.makedirs(raw_dir, exist_ok=True)
260
261    missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))]
262    if missing and not download:
263        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
264
265    for case in tqdm(missing, desc="Prepare Prostate CNB slides"):
266        image_name, mask_name = cases[case]
267        image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name))
268        for name, out_path in ((image_name, image_path), (mask_name, mask_path)):
269            if not os.path.exists(out_path):
270                _fetch_member(entries[name], out_path)
271
272        _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level)
273        os.remove(image_path)
274        os.remove(mask_path)
275
276    return preprocessed_dir

Download and preprocess the Prostate CNB dataset.

NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched. Use n_cases to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • n_cases: The number of slides to use, sorted by slide id ('B---
    -'). By default all 37 slides are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer. The preprocessed data is stored separately for each level.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the preprocessed data is stored.

def get_prostate_cnb_paths( path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False) -> List[str]:
279def get_prostate_cnb_paths(
280    path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False,
281) -> List[str]:
282    """Get paths to the Prostate CNB data.
283
284    Args:
285        path: Filepath to a folder where the data is downloaded for further processing.
286        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
287            By default all 37 slides are used.
288        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
289            micrometer.
290        download: Whether to download the data if it is not present.
291
292    Returns:
293        List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the
294        label data ('labels').
295    """
296    preprocessed_dir = get_prostate_cnb_data(path, n_cases, level, download)
297    case_ids = list(_get_cases(_read_zip_entries(path)))[:n_cases]
298    return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]

Get paths to the Prostate CNB data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • n_cases: The number of slides to use, sorted by slide id ('B---
    -'). By default all 37 slides are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the label data ('labels').

def get_prostate_cnb_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], n_cases: Optional[int] = None, level: int = 1, download: bool = False, label_dtype: torch.dtype = torch.int64, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
301def get_prostate_cnb_dataset(
302    path: Union[os.PathLike, str],
303    patch_shape: Tuple[int, int],
304    n_cases: Optional[int] = None,
305    level: int = 1,
306    download: bool = False,
307    label_dtype: torch.dtype = torch.int64,
308    resize_inputs: bool = False,
309    **kwargs
310) -> Dataset:
311    """Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images.
312
313    Args:
314        path: Filepath to a folder where the data is downloaded for further processing.
315        patch_shape: The patch shape to use for training.
316        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
317            By default all 37 slides are used.
318        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
319            micrometer.
320        download: Whether to download the data if it is not present.
321        label_dtype: The datatype of the labels.
322        resize_inputs: Whether to resize the input images.
323        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
324
325    Returns:
326        The segmentation dataset.
327    """
328    volume_paths = get_prostate_cnb_paths(path, n_cases, level, download)
329
330    if resize_inputs:
331        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
332        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
333            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
334        )
335
336    return torch_em.default_segmentation_dataset(
337        raw_paths=volume_paths,
338        raw_key="raw",
339        label_paths=volume_paths,
340        label_key="labels",
341        patch_shape=patch_shape,
342        label_dtype=label_dtype,
343        is_seg_dataset=True,
344        with_channels=True,
345        ndim=2,
346        **kwargs
347    )

Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • n_cases: The number of slides to use, sorted by slide id ('B---
    -'). By default all 37 slides are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
  • label_dtype: The datatype of the labels.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_prostate_cnb_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], n_cases: Optional[int] = None, level: int = 1, download: bool = False, label_dtype: torch.dtype = torch.int64, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
350def get_prostate_cnb_loader(
351    path: Union[os.PathLike, str],
352    batch_size: int,
353    patch_shape: Tuple[int, int],
354    n_cases: Optional[int] = None,
355    level: int = 1,
356    download: bool = False,
357    label_dtype: torch.dtype = torch.int64,
358    resize_inputs: bool = False,
359    **kwargs
360) -> DataLoader:
361    """Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images.
362
363    Args:
364        path: Filepath to a folder where the data is downloaded for further processing.
365        batch_size: The batch size for training.
366        patch_shape: The patch shape to use for training.
367        n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>').
368            By default all 37 slides are used.
369        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
370            micrometer.
371        download: Whether to download the data if it is not present.
372        label_dtype: The datatype of the labels.
373        resize_inputs: Whether to resize the input images.
374        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
375
376    Returns:
377        The DataLoader.
378    """
379    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
380    dataset = get_prostate_cnb_dataset(
381        path=path, patch_shape=patch_shape, n_cases=n_cases, level=level, download=download,
382        label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs
383    )
384    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • n_cases: The number of slides to use, sorted by slide id ('B---
    -'). By default all 37 slides are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
  • label_dtype: The datatype of the labels.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.