torch_em.data.datasets.histopathology.precise

The PRECISE dataset contains annotations for the semantic segmentation of prostate tissue and lesions in paired H&E and immunohistochemistry (IHC) whole-slide images of prostate core needle biopsies.

PRECISE (PRostate Expert-annotated Contiguous IHC-H&E Serial sEctions) consists of 54 slide pairs from 25 patients (sub-01 has three sessions, all other patients one). Each pair has a H&E and a HMWCK-AMACR (CKAPM + racemase) IHC whole-slide image, scanned at 0.243 micrometer per pixel with a 3DHISTECH Pannoramic scanner, and a pixel-level mask for each stain, drawn by expert uropathologists (24,387 annotations in total).

The label ids of the masks are:

  • 0: background
  • 1: tumor
  • 2: benign gland
  • 3: artifact
  • 4: high-grade prostatic intraepithelial neoplasia (HGPIN)
  • 5: intraductal carcinoma
  • 6: atypical intraductal proliferation
  • 7: stroma Not every class occurs in every slide.

The data is located at https://doi.org/10.5281/zenodo.20721779 as a single 55.6 GB zip archive with pyramidal OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole archive, this module reads its directory from the server and fetches only the requested slides and masks. Each slide is converted once into a chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the default level 1 is 0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use n_cases to only prepare a subset.

The publication for this dataset was not available at the time of implementation, so please cite the Zenodo record if you use it in your research.

  1"""The PRECISE dataset contains annotations for the semantic segmentation of prostate tissue and lesions
  2in paired H&E and immunohistochemistry (IHC) whole-slide images of prostate core needle biopsies.
  3
  4PRECISE (PRostate Expert-annotated Contiguous IHC-H&E Serial sEctions) consists of 54 slide pairs from 25 patients
  5(sub-01 has three sessions, all other patients one). Each pair has a H&E and a HMWCK-AMACR (CKAPM + racemase) IHC
  6whole-slide image, scanned at 0.243 micrometer per pixel with a 3DHISTECH Pannoramic scanner, and a pixel-level
  7mask for each stain, drawn by expert uropathologists (24,387 annotations in total).
  8
  9The label ids of the masks are:
 10- 0: background
 11- 1: tumor
 12- 2: benign gland
 13- 3: artifact
 14- 4: high-grade prostatic intraepithelial neoplasia (HGPIN)
 15- 5: intraductal carcinoma
 16- 6: atypical intraductal proliferation
 17- 7: stroma
 18Not every class occurs in every slide.
 19
 20The data is located at https://doi.org/10.5281/zenodo.20721779 as a single 55.6 GB zip archive with pyramidal
 21OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole archive, this module reads
 22its directory from the server and fetches only the requested slides and masks. Each slide is converted once into a
 23chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the default level 1 is
 240.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use `n_cases` to only prepare a subset.
 25
 26The publication for this dataset was not available at the time of implementation, so please cite the Zenodo record
 27if you use it in your research.
 28"""
 29
 30import os
 31import re
 32import json
 33import uuid
 34import zlib
 35import struct
 36from typing import List, Literal, Optional, Tuple, Union
 37
 38from tqdm import tqdm
 39
 40import torch
 41
 42from torch.utils.data import Dataset, DataLoader
 43
 44import torch_em
 45
 46from .. import util
 47
 48
 49RECORD_URL = "https://zenodo.org/api/records/20721779/files"
 50ZIP_URL = f"{RECORD_URL}/data.zip/content"
 51
 52SMALL_FILES = {
 53    "label_descriptions.json": "31f9687f2611fce5a56e5f983c9a97b6ae0f5d1b6171be65e9bbc961ad2ec11c",
 54    "participants.csv": "8b5f18120b83e84eb235eb746eb005f8dbbd0b9d46c30d248c90c0a7b31cdcaf",
 55}
 56
 57STAINS = {"he": "h-e", "ihc": "hmwck-amacr"}
 58
 59CLASS_NAMES = {
 60    0: "background",
 61    1: "tumor",
 62    2: "benign gland",
 63    3: "artifact",
 64    4: "high-grade prostatic intraepithelial neoplasia",
 65    5: "intraductal carcinoma",
 66    6: "atypical intraductal proliferation",
 67    7: "stroma",
 68}
 69
 70N_LEVELS = 6
 71
 72
 73def _range_get(url, start, end, **kwargs):
 74    import requests
 75
 76    response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, **kwargs)
 77    response.raise_for_status()
 78    return response
 79
 80
 81def _read_zip_entries(path):
 82    cache_path = os.path.join(path, "zip_entries.json")
 83    if os.path.exists(cache_path):
 84        with open(cache_path) as f:
 85            return json.load(f)
 86
 87    import requests
 88
 89    size = int(requests.head(ZIP_URL, allow_redirects=True).headers["Content-Length"])
 90    tail = _range_get(ZIP_URL, size - 65557, size - 1).content
 91    eocd = tail[tail.rfind(b"PK\x05\x06"):]
 92    n_entries, cd_size, cd_offset = struct.unpack("<HII", eocd[10:20])
 93    if cd_offset == 0xFFFFFFFF or n_entries == 0xFFFF:
 94        locator = tail[tail.rfind(b"PK\x06\x07"):][:20]
 95        zip64_offset = struct.unpack("<Q", locator[8:16])[0]
 96        record = _range_get(ZIP_URL, zip64_offset, zip64_offset + 55).content
 97        n_entries, cd_size, cd_offset = struct.unpack("<QQQ", record[32:56])
 98
 99    directory = _range_get(ZIP_URL, cd_offset, cd_offset + cd_size - 1).content
100    entries, pos = {}, 0
101    while pos < len(directory):
102        (signature, _, _, _, _, _, _, crc, comp_size, size, name_len, extra_len, comment_len, _, _, _,
103         header_offset) = struct.unpack("<IHHHHHHIIIHHHHHII", directory[pos:pos + 46])
104        assert signature == 0x02014B50, "The zip directory could not be parsed."
105        name = directory[pos + 46:pos + 46 + name_len].decode()
106        extra = directory[pos + 46 + name_len:pos + 46 + name_len + extra_len]
107        extra_pos = 0
108        while extra_pos < len(extra):
109            field_id, field_size = struct.unpack("<HH", extra[extra_pos:extra_pos + 4])
110            if field_id == 1:
111                field, k = extra[extra_pos + 4:extra_pos + 4 + field_size], 0
112                if size == 0xFFFFFFFF:
113                    size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8
114                if comp_size == 0xFFFFFFFF:
115                    comp_size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8
116                if header_offset == 0xFFFFFFFF:
117                    header_offset = struct.unpack("<Q", field[k:k + 8])[0]
118            extra_pos += 4 + field_size
119        if size > 0:
120            entries[name] = {"crc": crc, "comp_size": comp_size, "size": size, "offset": header_offset}
121        pos += 46 + name_len + extra_len + comment_len
122
123    tmp_path = f"{cache_path}.{uuid.uuid4().hex}.tmp"
124    with open(tmp_path, "w") as f:
125        json.dump(entries, f)
126    os.replace(tmp_path, cache_path)
127    return entries
128
129
130def _fetch_member(entry, out_path):
131    header = _range_get(ZIP_URL, entry["offset"], entry["offset"] + 29).content
132    name_len, extra_len = struct.unpack("<HH", header[26:30])
133    start = entry["offset"] + 30 + name_len + extra_len
134
135    decompressor, crc, n_bytes = zlib.decompressobj(-15), 0, 0
136    tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp"
137    with _range_get(ZIP_URL, start, start + entry["comp_size"] - 1, stream=True) as response, \
138            open(tmp_path, "wb") as f:
139        for chunk in response.iter_content(1 << 20):
140            data = decompressor.decompress(chunk)
141            crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data)
142            f.write(data)
143        data = decompressor.flush()
144        crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data)
145        f.write(data)
146
147    if crc != entry["crc"] or n_bytes != entry["size"]:
148        os.remove(tmp_path)
149        raise RuntimeError(f"The download of {os.path.basename(out_path)} is corrupted, please try again.")
150    os.replace(tmp_path, out_path)
151
152
153def _get_cases(entries, stain):
154    stain_dir = STAINS[stain]
155    stain = re.escape(stain_dir)
156    pattern = re.compile(rf"^data/sub-\d+/ses-\d+/wsi_{stain}/(sub-\d+_ses-\d+)_{stain}\.ome\.tif$")
157    cases = {}
158    for name in entries:
159        match = pattern.match(name)
160        if match:
161            case = match.group(1)
162            mask_name = name.replace(".ome.tif", "_mask.ome.tif")
163            assert mask_name in entries, f"Cannot find the mask for {name}."
164            cases[case] = (name, mask_name)
165    return dict(sorted(cases.items()))
166
167
168def _open_level(series, level_index):
169    import zarr
170
171    # The pyramidal TIFFs are tiled, so the zarr view reads only the requested tiles.
172    array = zarr.open(series.aszarr(), mode="r")
173    return array if hasattr(array, "shape") else array[str(level_index)]
174
175
176def _convert_case(image_path, mask_path, out_path, level, tile=4096):
177    import h5py
178    import tifffile
179    import numpy as np
180
181    image_series = tifffile.TiffFile(image_path).series[0]
182    mask_series = tifffile.TiffFile(mask_path).series[0]
183    image, mask = _open_level(image_series, level), _open_level(mask_series, level)
184
185    # The pyramid levels of the image and mask are rounded differently, so the mask level can be a few pixels off.
186    # It is mapped to the image grid by nearest neighbor sampling, unless the shapes really differ.
187    height, width = image.shape[:2]
188    if any(abs(image.shape[i] - mask.shape[i]) > 1e-3 * image.shape[i] + 1 for i in range(2)):
189        raise RuntimeError(f"The image {image.shape} and mask {mask.shape} shapes of {image_path} do not match.")
190    rows = np.minimum(((np.arange(height) + 0.5) * mask.shape[0] / height).astype(int), mask.shape[0] - 1)
191    cols = np.minimum(((np.arange(width) + 0.5) * mask.shape[1] / width).astype(int), mask.shape[1] - 1)
192
193    tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp"
194    with h5py.File(tmp_path, "w") as f:
195        raw = f.create_dataset(
196            "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(3, 512, 512)
197        )
198        labels = f.create_dataset(
199            "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512)
200        )
201        for y in range(0, height, tile):
202            for x in range(0, width, tile):
203                bb = (slice(y, min(y + tile, height)), slice(x, min(x + tile, width)))
204                raw[(slice(None),) + bb] = image[bb].transpose(2, 0, 1)
205                r, c = rows[bb[0]], cols[bb[1]]
206                mask_block = mask[r[0]:r[-1] + 1, c[0]:c[-1] + 1]
207                labels[bb] = mask_block[np.ix_(r - r[0], c - c[0])]
208    os.replace(tmp_path, out_path)
209
210
211def get_precise_data(
212    path: Union[os.PathLike, str],
213    stain: Literal["he", "ihc"] = "he",
214    n_cases: Optional[int] = None,
215    level: int = 1,
216    download: bool = False,
217) -> str:
218    """Download and preprocess the PRECISE dataset.
219
220    NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched.
221    Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download.
222
223    Args:
224        path: Filepath to a folder where the data is downloaded for further processing.
225        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
226        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
227            By default all 54 slides of the stain are used.
228        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
229            micrometer. The preprocessed data is stored separately for each level.
230        download: Whether to download the data if it is not present.
231
232    Returns:
233        Filepath to the folder where the preprocessed data is stored.
234    """
235    if stain not in STAINS:
236        raise ValueError(f"'{stain}' is not a valid stain. Choose one of {list(STAINS)}.")
237    if level not in range(N_LEVELS):
238        raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.")
239
240    os.makedirs(path, exist_ok=True)
241    for filename, checksum in SMALL_FILES.items():
242        util.download_source(
243            path=os.path.join(path, filename), url=f"{RECORD_URL}/{filename}/content", download=download,
244            checksum=checksum,
245        )
246
247    if not download and not os.path.exists(os.path.join(path, "zip_entries.json")):
248        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
249    entries = _read_zip_entries(path)
250    cases = _get_cases(entries, stain)
251    case_ids = list(cases)[:n_cases]
252
253    preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}", stain)
254    raw_dir = os.path.join(path, "raw", stain)
255    os.makedirs(preprocessed_dir, exist_ok=True)
256    os.makedirs(raw_dir, exist_ok=True)
257
258    missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))]
259    if missing and not download:
260        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
261
262    for case in tqdm(missing, desc="Prepare PRECISE slides"):
263        image_name, mask_name = cases[case]
264        image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name))
265        for name, out_path in ((image_name, image_path), (mask_name, mask_path)):
266            if not os.path.exists(out_path):
267                _fetch_member(entries[name], out_path)
268
269        _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level)
270        os.remove(image_path)
271        os.remove(mask_path)
272
273    return preprocessed_dir
274
275
276def get_precise_paths(
277    path: Union[os.PathLike, str],
278    stain: Literal["he", "ihc"] = "he",
279    n_cases: Optional[int] = None,
280    level: int = 1,
281    download: bool = False,
282) -> List[str]:
283    """Get paths to the PRECISE data.
284
285    Args:
286        path: Filepath to a folder where the data is downloaded for further processing.
287        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
288        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
289            By default all 54 slides of the stain are used.
290        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
291            micrometer.
292        download: Whether to download the data if it is not present.
293
294    Returns:
295        List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the
296        label data ('labels').
297    """
298    preprocessed_dir = get_precise_data(path, stain, n_cases, level, download)
299    case_ids = list(_get_cases(_read_zip_entries(path), stain))[:n_cases]
300    return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]
301
302
303def get_precise_dataset(
304    path: Union[os.PathLike, str],
305    patch_shape: Tuple[int, int],
306    stain: Literal["he", "ihc"] = "he",
307    n_cases: Optional[int] = None,
308    level: int = 1,
309    download: bool = False,
310    label_dtype: torch.dtype = torch.int64,
311    resize_inputs: bool = False,
312    **kwargs
313) -> Dataset:
314    """Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images.
315
316    Args:
317        path: Filepath to a folder where the data is downloaded for further processing.
318        patch_shape: The patch shape to use for training.
319        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
320        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
321            By default all 54 slides of the stain are used.
322        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
323            micrometer.
324        download: Whether to download the data if it is not present.
325        label_dtype: The datatype of the labels.
326        resize_inputs: Whether to resize the input images.
327        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
328
329    Returns:
330        The segmentation dataset.
331    """
332    volume_paths = get_precise_paths(path, stain, n_cases, level, download)
333
334    if resize_inputs:
335        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
336        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
337            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
338        )
339
340    return torch_em.default_segmentation_dataset(
341        raw_paths=volume_paths,
342        raw_key="raw",
343        label_paths=volume_paths,
344        label_key="labels",
345        patch_shape=patch_shape,
346        label_dtype=label_dtype,
347        is_seg_dataset=True,
348        with_channels=True,
349        ndim=2,
350        **kwargs
351    )
352
353
354def get_precise_loader(
355    path: Union[os.PathLike, str],
356    batch_size: int,
357    patch_shape: Tuple[int, int],
358    stain: Literal["he", "ihc"] = "he",
359    n_cases: Optional[int] = None,
360    level: int = 1,
361    download: bool = False,
362    label_dtype: torch.dtype = torch.int64,
363    resize_inputs: bool = False,
364    **kwargs
365) -> DataLoader:
366    """Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images.
367
368    Args:
369        path: Filepath to a folder where the data is downloaded for further processing.
370        batch_size: The batch size for training.
371        patch_shape: The patch shape to use for training.
372        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
373        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
374            By default all 54 slides of the stain are used.
375        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
376            micrometer.
377        download: Whether to download the data if it is not present.
378        label_dtype: The datatype of the labels.
379        resize_inputs: Whether to resize the input images.
380        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
381
382    Returns:
383        The DataLoader.
384    """
385    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
386    dataset = get_precise_dataset(
387        path=path, patch_shape=patch_shape, stain=stain, n_cases=n_cases, level=level, download=download,
388        label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs
389    )
390    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
RECORD_URL = 'https://zenodo.org/api/records/20721779/files'
ZIP_URL = 'https://zenodo.org/api/records/20721779/files/data.zip/content'
SMALL_FILES = {'label_descriptions.json': '31f9687f2611fce5a56e5f983c9a97b6ae0f5d1b6171be65e9bbc961ad2ec11c', 'participants.csv': '8b5f18120b83e84eb235eb746eb005f8dbbd0b9d46c30d248c90c0a7b31cdcaf'}
STAINS = {'he': 'h-e', 'ihc': 'hmwck-amacr'}
CLASS_NAMES = {0: 'background', 1: 'tumor', 2: 'benign gland', 3: 'artifact', 4: 'high-grade prostatic intraepithelial neoplasia', 5: 'intraductal carcinoma', 6: 'atypical intraductal proliferation', 7: 'stroma'}
N_LEVELS = 6
def get_precise_data( path: Union[os.PathLike, str], stain: Literal['he', 'ihc'] = 'he', n_cases: Optional[int] = None, level: int = 1, download: bool = False) -> str:
212def get_precise_data(
213    path: Union[os.PathLike, str],
214    stain: Literal["he", "ihc"] = "he",
215    n_cases: Optional[int] = None,
216    level: int = 1,
217    download: bool = False,
218) -> str:
219    """Download and preprocess the PRECISE dataset.
220
221    NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched.
222    Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download.
223
224    Args:
225        path: Filepath to a folder where the data is downloaded for further processing.
226        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
227        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
228            By default all 54 slides of the stain are used.
229        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
230            micrometer. The preprocessed data is stored separately for each level.
231        download: Whether to download the data if it is not present.
232
233    Returns:
234        Filepath to the folder where the preprocessed data is stored.
235    """
236    if stain not in STAINS:
237        raise ValueError(f"'{stain}' is not a valid stain. Choose one of {list(STAINS)}.")
238    if level not in range(N_LEVELS):
239        raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.")
240
241    os.makedirs(path, exist_ok=True)
242    for filename, checksum in SMALL_FILES.items():
243        util.download_source(
244            path=os.path.join(path, filename), url=f"{RECORD_URL}/{filename}/content", download=download,
245            checksum=checksum,
246        )
247
248    if not download and not os.path.exists(os.path.join(path, "zip_entries.json")):
249        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
250    entries = _read_zip_entries(path)
251    cases = _get_cases(entries, stain)
252    case_ids = list(cases)[:n_cases]
253
254    preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}", stain)
255    raw_dir = os.path.join(path, "raw", stain)
256    os.makedirs(preprocessed_dir, exist_ok=True)
257    os.makedirs(raw_dir, exist_ok=True)
258
259    missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))]
260    if missing and not download:
261        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
262
263    for case in tqdm(missing, desc="Prepare PRECISE slides"):
264        image_name, mask_name = cases[case]
265        image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name))
266        for name, out_path in ((image_name, image_path), (mask_name, mask_path)):
267            if not os.path.exists(out_path):
268                _fetch_member(entries[name], out_path)
269
270        _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level)
271        os.remove(image_path)
272        os.remove(mask_path)
273
274    return preprocessed_dir

Download and preprocess the PRECISE dataset.

NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched. Use n_cases to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
  • n_cases: The number of slides to use, sorted by slide id ('sub-_ses-'). By default all 54 slides of the stain are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer. The preprocessed data is stored separately for each level.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the preprocessed data is stored.

def get_precise_paths( path: Union[os.PathLike, str], stain: Literal['he', 'ihc'] = 'he', n_cases: Optional[int] = None, level: int = 1, download: bool = False) -> List[str]:
277def get_precise_paths(
278    path: Union[os.PathLike, str],
279    stain: Literal["he", "ihc"] = "he",
280    n_cases: Optional[int] = None,
281    level: int = 1,
282    download: bool = False,
283) -> List[str]:
284    """Get paths to the PRECISE data.
285
286    Args:
287        path: Filepath to a folder where the data is downloaded for further processing.
288        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
289        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
290            By default all 54 slides of the stain are used.
291        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
292            micrometer.
293        download: Whether to download the data if it is not present.
294
295    Returns:
296        List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the
297        label data ('labels').
298    """
299    preprocessed_dir = get_precise_data(path, stain, n_cases, level, download)
300    case_ids = list(_get_cases(_read_zip_entries(path), stain))[:n_cases]
301    return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]

Get paths to the PRECISE data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
  • n_cases: The number of slides to use, sorted by slide id ('sub-_ses-'). By default all 54 slides of the stain are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the label data ('labels').

def get_precise_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], stain: Literal['he', 'ihc'] = 'he', n_cases: Optional[int] = None, level: int = 1, download: bool = False, label_dtype: torch.dtype = torch.int64, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
304def get_precise_dataset(
305    path: Union[os.PathLike, str],
306    patch_shape: Tuple[int, int],
307    stain: Literal["he", "ihc"] = "he",
308    n_cases: Optional[int] = None,
309    level: int = 1,
310    download: bool = False,
311    label_dtype: torch.dtype = torch.int64,
312    resize_inputs: bool = False,
313    **kwargs
314) -> Dataset:
315    """Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images.
316
317    Args:
318        path: Filepath to a folder where the data is downloaded for further processing.
319        patch_shape: The patch shape to use for training.
320        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
321        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
322            By default all 54 slides of the stain are used.
323        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
324            micrometer.
325        download: Whether to download the data if it is not present.
326        label_dtype: The datatype of the labels.
327        resize_inputs: Whether to resize the input images.
328        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
329
330    Returns:
331        The segmentation dataset.
332    """
333    volume_paths = get_precise_paths(path, stain, n_cases, level, download)
334
335    if resize_inputs:
336        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
337        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
338            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
339        )
340
341    return torch_em.default_segmentation_dataset(
342        raw_paths=volume_paths,
343        raw_key="raw",
344        label_paths=volume_paths,
345        label_key="labels",
346        patch_shape=patch_shape,
347        label_dtype=label_dtype,
348        is_seg_dataset=True,
349        with_channels=True,
350        ndim=2,
351        **kwargs
352    )

Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
  • n_cases: The number of slides to use, sorted by slide id ('sub-_ses-'). By default all 54 slides of the stain are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
  • label_dtype: The datatype of the labels.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_precise_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], stain: Literal['he', 'ihc'] = 'he', n_cases: Optional[int] = None, level: int = 1, download: bool = False, label_dtype: torch.dtype = torch.int64, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
355def get_precise_loader(
356    path: Union[os.PathLike, str],
357    batch_size: int,
358    patch_shape: Tuple[int, int],
359    stain: Literal["he", "ihc"] = "he",
360    n_cases: Optional[int] = None,
361    level: int = 1,
362    download: bool = False,
363    label_dtype: torch.dtype = torch.int64,
364    resize_inputs: bool = False,
365    **kwargs
366) -> DataLoader:
367    """Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images.
368
369    Args:
370        path: Filepath to a folder where the data is downloaded for further processing.
371        batch_size: The batch size for training.
372        patch_shape: The patch shape to use for training.
373        stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
374        n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>').
375            By default all 54 slides of the stain are used.
376        level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level
377            micrometer.
378        download: Whether to download the data if it is not present.
379        label_dtype: The datatype of the labels.
380        resize_inputs: Whether to resize the input images.
381        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
382
383    Returns:
384        The DataLoader.
385    """
386    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
387    dataset = get_precise_dataset(
388        path=path, patch_shape=patch_shape, stain=stain, n_cases=n_cases, level=level, download=download,
389        label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs
390    )
391    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
  • n_cases: The number of slides to use, sorted by slide id ('sub-_ses-'). By default all 54 slides of the stain are used.
  • level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
  • download: Whether to download the data if it is not present.
  • label_dtype: The datatype of the labels.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.