torch_em.data.datasets.light_microscopy.xenium_lung_treg

This dataset contains cell and nucleus segmentation masks for whole-slide fluorescence images of mouse lung tissue, from the BioImage Archive study S-BIAD3880.

The study images FFPE sections of adult female FVB/N mouse lung with the 10x Genomics Xenium Prime 5K platform, to map the preinvasive regulatory T cell niches that form around NTCU-induced airway lesions and how the PI3K-delta inhibitor PI-3065 changes them. It holds 8 samples: 2 untreated controls, 3 NTCU-treated mice, and 3 NTCU-treated mice that also got PI-3065. See TREATMENT for which group each sample belongs to.

Every sample pairs a set of morphology images with the masks that the Xenium Onboard Analysis (XOA) pipeline produced. Channel 0 always holds the DAPI nuclear stain. The other channels hold the boundary and interior stains of the multi-tissue stain mix, which XOA uses to draw the cell outlines.

NOTE: The masks are the output of XOA, not a manual annotation. XOA segments the nuclei on DAPI, then grows each cell from its nucleus using the boundary and interior stains. Where no stain supports a boundary it falls back to expanding the nucleus by 5 micrometer, so a share of the cells are dilated nuclei rather than observed outlines. The 'nuc_expansion_fraction' attribute of each h5 file reports that share. Treat this data as a weak label source for pretraining, not as a benchmark.

NOTE: A cell can hold more than one nucleus, so the two label sets do not match one to one and their ids do not correspond.

NOTE: This dataset is distinct from the Xenium samples in torch_em.data.datasets.light_microscopy.xenium, which are hosted directly by 10x Genomics and cover human and mouse tissue other than lung.

NOTE: The images are whole slides that hold large empty regions. Pass a sampler such as torch_em.data.MinInstanceSampler to avoid drawing empty patches.

The data is hosted at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3880 under the CC0 1.0 license. Please cite the corresponding study page if you use this data in your research.

  1"""This dataset contains cell and nucleus segmentation masks for whole-slide fluorescence
  2images of mouse lung tissue, from the BioImage Archive study S-BIAD3880.
  3
  4The study images FFPE sections of adult female FVB/N mouse lung with the 10x Genomics
  5Xenium Prime 5K platform, to map the preinvasive regulatory T cell niches that form around
  6NTCU-induced airway lesions and how the PI3K-delta inhibitor PI-3065 changes them. It holds
  78 samples: 2 untreated controls, 3 NTCU-treated mice, and 3 NTCU-treated mice that also got
  8PI-3065. See `TREATMENT` for which group each sample belongs to.
  9
 10Every sample pairs a set of morphology images with the masks that the Xenium Onboard
 11Analysis (XOA) pipeline produced. Channel 0 always holds the DAPI nuclear stain. The other
 12channels hold the boundary and interior stains of the multi-tissue stain mix, which XOA uses
 13to draw the cell outlines.
 14
 15NOTE: The masks are the output of XOA, not a manual annotation. XOA segments the nuclei on
 16DAPI, then grows each cell from its nucleus using the boundary and interior stains. Where no
 17stain supports a boundary it falls back to expanding the nucleus by 5 micrometer, so a share
 18of the cells are dilated nuclei rather than observed outlines. The 'nuc_expansion_fraction'
 19attribute of each h5 file reports that share. Treat this data as a weak label source for
 20pretraining, not as a benchmark.
 21
 22NOTE: A cell can hold more than one nucleus, so the two label sets do not match one to one
 23and their ids do not correspond.
 24
 25NOTE: This dataset is distinct from the Xenium samples in `torch_em.data.datasets.light_microscopy.xenium`,
 26which are hosted directly by 10x Genomics and cover human and mouse tissue other than lung.
 27
 28NOTE: The images are whole slides that hold large empty regions. Pass a sampler such as
 29`torch_em.data.MinInstanceSampler` to avoid drawing empty patches.
 30
 31The data is hosted at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3880 under the
 32CC0 1.0 license. Please cite the corresponding study page if you use this data in your research.
 33"""
 34
 35import os
 36import json
 37import zlib
 38import shutil
 39import struct
 40import zipfile
 41from glob import glob
 42from natsort import natsorted
 43from contextlib import contextmanager
 44from typing import Dict, List, Literal, Optional, Sequence, Tuple, Union
 45
 46import requests
 47
 48from torch.utils.data import DataLoader, Dataset
 49
 50import torch_em
 51
 52from .. import util
 53
 54
 55BASE_URL = "https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload"
 56
 57URLS = {
 58    "ctrl69": f"{BASE_URL}/ctrl69_xenium_output.zip",
 59    "ctrl70": f"{BASE_URL}/ctrl70_xenium_output.zip",
 60    "ntcu06": f"{BASE_URL}/ntcu06_xenium_output.zip",
 61    "ntcu36": f"{BASE_URL}/ntcu36_xenium_output.zip",
 62    "ntcu46": f"{BASE_URL}/ntcu46_xenium_output.zip",
 63    "pi42": f"{BASE_URL}/pi42_xenium_output.zip",
 64    "pi51": f"{BASE_URL}/pi51_xenium_output.zip",
 65    "pi57": f"{BASE_URL}/pi57_xenium_output.zip",
 66}
 67
 68# The treatment group of every sample, for the h5 attributes.
 69TREATMENT = {
 70    "ctrl69": "control",
 71    "ctrl70": "control",
 72    "ntcu06": "ntcu",
 73    "ntcu36": "ntcu",
 74    "ntcu46": "ntcu",
 75    "pi42": "ntcu_pi3065",
 76    "pi51": "ntcu_pi3065",
 77    "pi57": "ntcu_pi3065",
 78}
 79
 80LABEL_CHANNELS = {"nuclei": "0", "cells": "1"}
 81
 82BLOCK = 4096
 83
 84
 85def _read_range(url: str, start: int, end: int) -> bytes:
 86    """Read the bytes from `start` to `end`, both included."""
 87    response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, stream=True)
 88    response.raise_for_status()
 89    return response.content
 90
 91
 92def _read_zip64_extra(extra: bytes, uncompressed: int, compressed: int, offset: int) -> Tuple[int, int, int]:
 93    """Replace the fields that overflowed the 32 bit form with their zip64 values."""
 94    position = 0
 95    while position + 4 <= len(extra):
 96        tag, size = struct.unpack("<HH", extra[position:position + 4])
 97        if tag == 0x0001:
 98            values = list(struct.unpack(f"<{size // 8}Q", extra[position + 4:position + 4 + size]))
 99            if uncompressed == 0xFFFFFFFF:
100                uncompressed = values.pop(0)
101            if compressed == 0xFFFFFFFF:
102                compressed = values.pop(0)
103            if offset == 0xFFFFFFFF:
104                offset = values.pop(0)
105            break
106        position += 4 + size
107    return uncompressed, compressed, offset
108
109
110def _remote_zip_index(url: str) -> Dict[str, Dict[str, int]]:
111    """Read the central directory of a remote zip, which lists every member and where it sits.
112
113    This turns the bundle into a set of files that can be read one by one, so that a download
114    can skip the transcripts, the per-cycle QC images, and the expression matrix, which
115    segmentation never touches.
116    """
117    response = requests.head(url)
118    response.raise_for_status()
119    size = int(response.headers["Content-Length"])
120
121    tail = _read_range(url, max(0, size - 262144), size - 1)
122    position = tail.rfind(b"PK\x06\x06")
123    if position == -1:
124        position = tail.rfind(b"PK\x05\x06")
125        if position == -1:
126            raise RuntimeError(f"Could not find the central directory of the zip at {url}.")
127        directory_size, directory_offset = struct.unpack("<II", tail[position + 12:position + 20])
128    else:
129        directory_size, directory_offset = struct.unpack("<QQ", tail[position + 40:position + 56])
130
131    directory = _read_range(url, directory_offset, directory_offset + directory_size + 128)
132    index, position = {}, 0
133    while position + 46 <= len(directory) and directory[position:position + 4] == b"PK\x01\x02":
134        method = struct.unpack("<H", directory[position + 10:position + 12])[0]
135        crc = struct.unpack("<I", directory[position + 16:position + 20])[0]
136        compressed, uncompressed = struct.unpack("<II", directory[position + 20:position + 28])
137        name_length, extra_length, comment_length = struct.unpack("<HHH", directory[position + 28:position + 34])
138        offset = struct.unpack("<I", directory[position + 42:position + 46])[0]
139        name = directory[position + 46:position + 46 + name_length].decode("utf-8", errors="replace")
140        extra = directory[position + 46 + name_length:position + 46 + name_length + extra_length]
141        uncompressed, compressed, offset = _read_zip64_extra(extra, uncompressed, compressed, offset)
142        index[name] = {"method": method, "crc": crc, "compressed": compressed, "offset": offset}
143        position += 46 + name_length + extra_length + comment_length
144    return index
145
146
147def _find_root(index: Dict[str, Dict[str, int]], sample: str) -> str:
148    """Find the single top-level output folder that every member of the bundle sits under."""
149    roots = {name.split("/")[0] for name in index if "/" in name}
150    if len(roots) != 1:
151        raise RuntimeError(f"Expected exactly one output folder in the bundle for '{sample}', found {roots}.")
152    return next(iter(roots)) + "/"
153
154
155def _download_member(url: str, name: str, entry: Dict[str, int], output_path: str) -> None:
156    """Read one member out of a remote zip and write it, checking its crc32."""
157    from tqdm import tqdm
158
159    # The local header repeats the name and carries its own extra field, so the data of the
160    # member starts behind a header whose length only the header itself gives.
161    header = _read_range(url, entry["offset"], entry["offset"] + 29)
162    if header[:4] != b"PK\x03\x04":
163        raise RuntimeError(f"The member '{name}' does not start with a local header.")
164    name_length, extra_length = struct.unpack("<HH", header[26:30])
165    start = entry["offset"] + 30 + name_length + extra_length
166
167    if entry["method"] not in (0, 8):
168        raise RuntimeError(f"The member '{name}' uses the unsupported compression method {entry['method']}.")
169    decompressor = zlib.decompressobj(-zlib.MAX_WBITS) if entry["method"] == 8 else None
170
171    crc = 0
172    temporary_path = f"{output_path}.incomplete"
173    os.makedirs(os.path.dirname(output_path), exist_ok=True)
174
175    headers = {"Range": f"bytes={start}-{start + entry['compressed'] - 1}"}
176    with requests.get(url, headers=headers, stream=True) as response:
177        response.raise_for_status()
178        description = f"Download {name}"
179        with open(temporary_path, "wb") as f:
180            with tqdm(total=entry["compressed"], unit="B", unit_scale=True, desc=description) as progress:
181                for chunk in response.iter_content(chunk_size=1 << 20):
182                    progress.update(len(chunk))
183                    block = decompressor.decompress(chunk) if decompressor is not None else chunk
184                    crc = zlib.crc32(block, crc)
185                    f.write(block)
186            if decompressor is not None:
187                block = decompressor.flush()
188                crc = zlib.crc32(block, crc)
189                f.write(block)
190
191    if crc != entry["crc"]:
192        os.remove(temporary_path)
193        raise RuntimeError(f"The member '{name}' has the crc32 {crc}, but the zip lists {entry['crc']}.")
194    os.replace(temporary_path, output_path)
195
196
197def _download_bundle(url: str, bundle_dir: str, sample: str, download: bool, with_stains: bool) -> str:
198    """Download the members that this loader reads, and skip the rest of the bundle."""
199    if os.path.exists(bundle_dir):
200        return bundle_dir
201    if not download:
202        raise RuntimeError(f"Cannot find the data at {bundle_dir}, but download was set to False")
203
204    index = _remote_zip_index(url)
205    root = _find_root(index, sample)
206    channels = natsorted(
207        n for n in index if n.startswith(f"{root}morphology_focus/") and n.endswith(".ome.tif")
208    )
209    if not channels:
210        raise RuntimeError(f"The bundle for '{sample}' holds no morphology channel.")
211    if not with_stains:
212        channels = channels[:1]
213
214    names = [f"{root}experiment.xenium", f"{root}cells.zarr.zip"] + channels
215    missing = [n for n in names if n not in index]
216    if missing:
217        raise RuntimeError(f"The bundle for '{sample}' is missing {missing}.")
218
219    temporary_dir = f"{bundle_dir}.tmp"
220    if os.path.exists(temporary_dir):
221        shutil.rmtree(temporary_dir)
222    for name in names:
223        relative_name = name[len(root):]
224        _download_member(url, name, index[name], os.path.join(temporary_dir, relative_name))
225
226    os.replace(temporary_dir, bundle_dir)
227    return bundle_dir
228
229
230def _channel_paths(bundle_dir: str) -> List[str]:
231    """Return the morphology channels, DAPI first."""
232    paths = natsorted(glob(os.path.join(bundle_dir, "morphology_focus", "*.ome.tif")))
233    if not paths:
234        raise RuntimeError(f"Could not find any morphology channel in {bundle_dir}.")
235    return paths
236
237
238@contextmanager
239def _open_image(path: str):
240    """Open one morphology channel at full resolution, without reading it into memory.
241
242    The channels form a multi file OME pyramid, which the OME reader of tifffile rejects.
243    Reading the file on its own gives the pyramid of this channel alone.
244    """
245    import zarr
246    import tifffile
247
248    with tifffile.TiffFile(path, is_ome=False) as tif:
249        store = tif.series[0].levels[0].aszarr()
250        try:
251            yield zarr.open(store, mode="r")["0"]
252        finally:
253            store.close()
254
255
256def _copy_blockwise(source, target, channel: Optional[int] = None) -> int:
257    """Copy a large 2d array block by block, to keep the memory use bounded.
258
259    The target must be indexed in one step, because indexing an h5 dataset twice writes into
260    a copy of the data rather than into the file.
261
262    Returns the largest value that was copied, which gives the instance count of a mask
263    without a second pass over the whole array.
264    """
265    height, width = source.shape[-2:]
266    largest = 0
267    for y in range(0, height, BLOCK):
268        for x in range(0, width, BLOCK):
269            box = (slice(y, min(y + BLOCK, height)), slice(x, min(x + BLOCK, width)))
270            block = source[box]
271            if channel is None:
272                target[box] = block
273            else:
274                target[(channel,) + box] = block
275            largest = max(largest, int(block.max()))
276    return largest
277
278
279def _create_h5(bundle_dir: str, output_path: str, sample: str, with_stains: bool) -> str:
280    import h5py
281    import zarr
282    from tqdm import tqdm
283
284    if os.path.exists(output_path):
285        return output_path
286
287    with open(os.path.join(bundle_dir, "experiment.xenium")) as f:
288        manifest = json.load(f)
289
290    masks_zip = os.path.join(bundle_dir, "cells.zarr.zip")
291    masks_dir = os.path.join(bundle_dir, "cells.zarr")
292    if not os.path.exists(masks_dir):
293        with zipfile.ZipFile(masks_zip) as archive:
294            archive.extractall(masks_dir)
295    masks = zarr.open(masks_dir, mode="r")["masks"]
296
297    channel_paths = _channel_paths(bundle_dir)
298    with _open_image(channel_paths[0]) as dapi:
299        shape = dapi.shape
300    for name, key in LABEL_CHANNELS.items():
301        if masks[key].shape != shape:
302            raise RuntimeError(
303                f"The {name} mask of '{sample}' has the shape {masks[key].shape}, "
304                f"but its DAPI image has the shape {shape}."
305            )
306
307    temporary_path = f"{output_path}.tmp"
308    with h5py.File(temporary_path, "w") as f:
309        f.attrs["sample"] = sample
310        f.attrs["treatment_group"] = TREATMENT[sample]
311        f.attrs["axes"] = "yx"
312        f.attrs["pixel_size"] = manifest["pixel_size"]
313        f.attrs["xoa_version"] = manifest["analysis_sw_version"]
314        f.attrs["segmentation_stain"] = manifest.get("segmentation_stain", "none")
315        f.attrs["nuc_expansion_fraction"] = manifest.get("segmented_cell_nuc_expansion_frac", float("nan"))
316        f.attrs["num_cells"] = manifest["num_cells"]
317        f.attrs["channel_files"] = [os.path.basename(p) for p in channel_paths]
318
319        chunks = (min(1024, shape[0]), min(1024, shape[1]))
320        raw = f.create_dataset("raw/dapi", shape=shape, dtype="uint16", chunks=chunks, compression="gzip")
321        raw.attrs["pixel_size"] = manifest["pixel_size"]
322        with _open_image(channel_paths[0]) as dapi:
323            _copy_blockwise(dapi, raw)
324
325        if with_stains:
326            stack_shape = (len(channel_paths),) + shape
327            stack = f.create_dataset(
328                "raw/stack", shape=stack_shape, dtype="uint16", chunks=(1,) + chunks, compression="gzip"
329            )
330            stack.attrs["pixel_size"] = manifest["pixel_size"]
331            for channel, channel_path in enumerate(tqdm(channel_paths, desc=f"Copy channels of '{sample}'")):
332                with _open_image(channel_path) as image:
333                    if image.shape != shape:
334                        raise RuntimeError(
335                            f"The channel {os.path.basename(channel_path)} of '{sample}' has the shape "
336                            f"{image.shape}, but its DAPI image has the shape {shape}."
337                        )
338                    _copy_blockwise(image, stack, channel=channel)
339
340        for name, key in LABEL_CHANNELS.items():
341            labels = f.create_dataset(
342                f"labels/{name}", shape=shape, dtype="uint32", chunks=chunks, compression="gzip"
343            )
344            labels.attrs["pixel_size"] = manifest["pixel_size"]
345            labels.attrs["num_instances"] = _copy_blockwise(masks[key], labels)
346
347    os.replace(temporary_path, output_path)
348    shutil.rmtree(masks_dir, ignore_errors=True)
349    return output_path
350
351
352def get_xenium_lung_treg_data(
353    path: Union[os.PathLike, str],
354    sample: str,
355    download: bool = False,
356    with_stains: bool = True,
357) -> str:
358    """Download one sample of the mouse lung Treg Xenium dataset.
359
360    This function reads the central directory of the remote zip that BioImage Archive
361    hosts for the sample, then reads only the members it needs (the manifest, the cell and
362    nucleus masks, and the morphology channels). It does not fetch the transcripts, the
363    per-cycle QC images, or the expression matrix, which cuts the download by a large margin.
364
365    Args:
366        path: Filepath to a folder where the downloaded data will be saved.
367        sample: The sample to use. See `URLS` for the available samples.
368        download: Whether to download the data if it is not present.
369        with_stains: Whether to also store the boundary and interior stains next to DAPI.
370            The cell masks rest on those stains, so a model that predicts cells needs them.
371
372    Returns:
373        The filepath to the h5 file of the sample.
374    """
375    if sample not in URLS:
376        raise ValueError(f"'{sample}' is not a valid sample. Choose from {list(URLS)}.")
377
378    output_path = os.path.join(path, "preprocessed", f"{sample}.h5")
379    if os.path.exists(output_path):
380        return output_path
381
382    os.makedirs(os.path.join(path, "preprocessed"), exist_ok=True)
383    bundle_dir = _download_bundle(URLS[sample], os.path.join(path, "bundles", sample), sample, download, with_stains)
384    _create_h5(bundle_dir, output_path, sample, with_stains)
385    shutil.rmtree(bundle_dir, ignore_errors=True)
386
387    return output_path
388
389
390def get_xenium_lung_treg_paths(
391    path: Union[os.PathLike, str],
392    sample: Optional[Union[str, Sequence[str]]] = None,
393    download: bool = False,
394    with_stains: bool = True,
395) -> List[str]:
396    """Get paths to the mouse lung Treg Xenium data.
397
398    Args:
399        path: Filepath to a folder where the downloaded data will be saved.
400        sample: The sample or samples to use. Defaults to all of them.
401        download: Whether to download the data if it is not present.
402        with_stains: Whether to also store the boundary and interior stains next to DAPI.
403
404    Returns:
405        List of filepaths for the h5 data.
406    """
407    if sample is None:
408        samples = list(URLS)
409    else:
410        samples = [sample] if isinstance(sample, str) else list(sample)
411
412    return [get_xenium_lung_treg_data(path, name, download, with_stains) for name in samples]
413
414
415def get_xenium_lung_treg_dataset(
416    path: Union[os.PathLike, str],
417    patch_shape: Tuple[int, int],
418    sample: Optional[Union[str, Sequence[str]]] = None,
419    label_channel: Literal["nuclei", "cells"] = "nuclei",
420    raw_channel: Literal["dapi", "stack"] = "dapi",
421    offsets: Optional[List[List[int]]] = None,
422    boundaries: bool = False,
423    binary: bool = False,
424    download: bool = False,
425    **kwargs,
426) -> Dataset:
427    """Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation.
428
429    Args:
430        path: Filepath to a folder where the downloaded data will be saved.
431        patch_shape: The 2D patch shape to use for training.
432        sample: The sample or samples to use. Defaults to all of them.
433        label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
434        raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone,
435            or 'stack' for all morphology channels.
436        offsets: Offset values for affinity computation used as target.
437        boundaries: Whether to compute boundaries as the target.
438        binary: Whether to use a binary segmentation target.
439        download: Whether to download the data if it is not present.
440        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
441
442    Returns:
443        The segmentation dataset.
444    """
445    if len(patch_shape) != 2:
446        raise ValueError(f"The Xenium patch shape must be two-dimensional, got {patch_shape}.")
447    if label_channel not in LABEL_CHANNELS:
448        raise ValueError(f"'{label_channel}' is not a valid label channel. Choose from {list(LABEL_CHANNELS)}.")
449    if raw_channel not in ("dapi", "stack"):
450        raise ValueError(f"'{raw_channel}' is not a valid raw channel. Choose from ['dapi', 'stack'].")
451
452    volume_paths = get_xenium_lung_treg_paths(path, sample, download, with_stains=raw_channel == "stack")
453
454    kwargs = util.update_kwargs(kwargs, "with_channels", raw_channel == "stack")
455    kwargs, _ = util.add_instance_label_transform(
456        kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary,
457    )
458    kwargs = util.ensure_transforms(ndim=2, **kwargs)
459
460    return torch_em.default_segmentation_dataset(
461        raw_paths=volume_paths,
462        raw_key=f"raw/{raw_channel}",
463        label_paths=volume_paths,
464        label_key=f"labels/{label_channel}",
465        patch_shape=patch_shape,
466        ndim=2,
467        **kwargs,
468    )
469
470
471def get_xenium_lung_treg_loader(
472    path: Union[os.PathLike, str],
473    batch_size: int,
474    patch_shape: Tuple[int, int],
475    sample: Optional[Union[str, Sequence[str]]] = None,
476    label_channel: Literal["nuclei", "cells"] = "nuclei",
477    raw_channel: Literal["dapi", "stack"] = "dapi",
478    offsets: Optional[List[List[int]]] = None,
479    boundaries: bool = False,
480    binary: bool = False,
481    download: bool = False,
482    **kwargs,
483) -> DataLoader:
484    """Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation.
485
486    Args:
487        path: Filepath to a folder where the downloaded data will be saved.
488        batch_size: The batch size for training.
489        patch_shape: The 2D patch shape to use for training.
490        sample: The sample or samples to use. Defaults to all of them.
491        label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
492        raw_channel: The images to use as input. Either 'dapi' or 'stack'.
493        offsets: Offset values for affinity computation used as target.
494        boundaries: Whether to compute boundaries as the target.
495        binary: Whether to use a binary segmentation target.
496        download: Whether to download the data if it is not present.
497        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader.
498
499    Returns:
500        The DataLoader.
501    """
502    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
503    dataset = get_xenium_lung_treg_dataset(
504        path=path,
505        patch_shape=patch_shape,
506        sample=sample,
507        label_channel=label_channel,
508        raw_channel=raw_channel,
509        offsets=offsets,
510        boundaries=boundaries,
511        binary=binary,
512        download=download,
513        **ds_kwargs,
514    )
515    return torch_em.get_data_loader(dataset, batch_size=batch_size, **loader_kwargs)
BASE_URL = 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload'
URLS = {'ctrl69': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/ctrl69_xenium_output.zip', 'ctrl70': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/ctrl70_xenium_output.zip', 'ntcu06': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/ntcu06_xenium_output.zip', 'ntcu36': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/ntcu36_xenium_output.zip', 'ntcu46': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/ntcu46_xenium_output.zip', 'pi42': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/pi42_xenium_output.zip', 'pi51': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/pi51_xenium_output.zip', 'pi57': 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload/pi57_xenium_output.zip'}
TREATMENT = {'ctrl69': 'control', 'ctrl70': 'control', 'ntcu06': 'ntcu', 'ntcu36': 'ntcu', 'ntcu46': 'ntcu', 'pi42': 'ntcu_pi3065', 'pi51': 'ntcu_pi3065', 'pi57': 'ntcu_pi3065'}
LABEL_CHANNELS = {'nuclei': '0', 'cells': '1'}
BLOCK = 4096
def get_xenium_lung_treg_data( path: Union[os.PathLike, str], sample: str, download: bool = False, with_stains: bool = True) -> str:
353def get_xenium_lung_treg_data(
354    path: Union[os.PathLike, str],
355    sample: str,
356    download: bool = False,
357    with_stains: bool = True,
358) -> str:
359    """Download one sample of the mouse lung Treg Xenium dataset.
360
361    This function reads the central directory of the remote zip that BioImage Archive
362    hosts for the sample, then reads only the members it needs (the manifest, the cell and
363    nucleus masks, and the morphology channels). It does not fetch the transcripts, the
364    per-cycle QC images, or the expression matrix, which cuts the download by a large margin.
365
366    Args:
367        path: Filepath to a folder where the downloaded data will be saved.
368        sample: The sample to use. See `URLS` for the available samples.
369        download: Whether to download the data if it is not present.
370        with_stains: Whether to also store the boundary and interior stains next to DAPI.
371            The cell masks rest on those stains, so a model that predicts cells needs them.
372
373    Returns:
374        The filepath to the h5 file of the sample.
375    """
376    if sample not in URLS:
377        raise ValueError(f"'{sample}' is not a valid sample. Choose from {list(URLS)}.")
378
379    output_path = os.path.join(path, "preprocessed", f"{sample}.h5")
380    if os.path.exists(output_path):
381        return output_path
382
383    os.makedirs(os.path.join(path, "preprocessed"), exist_ok=True)
384    bundle_dir = _download_bundle(URLS[sample], os.path.join(path, "bundles", sample), sample, download, with_stains)
385    _create_h5(bundle_dir, output_path, sample, with_stains)
386    shutil.rmtree(bundle_dir, ignore_errors=True)
387
388    return output_path

Download one sample of the mouse lung Treg Xenium dataset.

This function reads the central directory of the remote zip that BioImage Archive hosts for the sample, then reads only the members it needs (the manifest, the cell and nucleus masks, and the morphology channels). It does not fetch the transcripts, the per-cycle QC images, or the expression matrix, which cuts the download by a large margin.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • sample: The sample to use. See URLS for the available samples.
  • download: Whether to download the data if it is not present.
  • with_stains: Whether to also store the boundary and interior stains next to DAPI. The cell masks rest on those stains, so a model that predicts cells needs them.
Returns:

The filepath to the h5 file of the sample.

def get_xenium_lung_treg_paths( path: Union[os.PathLike, str], sample: Union[str, Sequence[str], NoneType] = None, download: bool = False, with_stains: bool = True) -> List[str]:
391def get_xenium_lung_treg_paths(
392    path: Union[os.PathLike, str],
393    sample: Optional[Union[str, Sequence[str]]] = None,
394    download: bool = False,
395    with_stains: bool = True,
396) -> List[str]:
397    """Get paths to the mouse lung Treg Xenium data.
398
399    Args:
400        path: Filepath to a folder where the downloaded data will be saved.
401        sample: The sample or samples to use. Defaults to all of them.
402        download: Whether to download the data if it is not present.
403        with_stains: Whether to also store the boundary and interior stains next to DAPI.
404
405    Returns:
406        List of filepaths for the h5 data.
407    """
408    if sample is None:
409        samples = list(URLS)
410    else:
411        samples = [sample] if isinstance(sample, str) else list(sample)
412
413    return [get_xenium_lung_treg_data(path, name, download, with_stains) for name in samples]

Get paths to the mouse lung Treg Xenium data.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • sample: The sample or samples to use. Defaults to all of them.
  • download: Whether to download the data if it is not present.
  • with_stains: Whether to also store the boundary and interior stains next to DAPI.
Returns:

List of filepaths for the h5 data.

def get_xenium_lung_treg_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], sample: Union[str, Sequence[str], NoneType] = None, label_channel: Literal['nuclei', 'cells'] = 'nuclei', raw_channel: Literal['dapi', 'stack'] = 'dapi', offsets: Optional[List[List[int]]] = None, boundaries: bool = False, binary: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
416def get_xenium_lung_treg_dataset(
417    path: Union[os.PathLike, str],
418    patch_shape: Tuple[int, int],
419    sample: Optional[Union[str, Sequence[str]]] = None,
420    label_channel: Literal["nuclei", "cells"] = "nuclei",
421    raw_channel: Literal["dapi", "stack"] = "dapi",
422    offsets: Optional[List[List[int]]] = None,
423    boundaries: bool = False,
424    binary: bool = False,
425    download: bool = False,
426    **kwargs,
427) -> Dataset:
428    """Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation.
429
430    Args:
431        path: Filepath to a folder where the downloaded data will be saved.
432        patch_shape: The 2D patch shape to use for training.
433        sample: The sample or samples to use. Defaults to all of them.
434        label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
435        raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone,
436            or 'stack' for all morphology channels.
437        offsets: Offset values for affinity computation used as target.
438        boundaries: Whether to compute boundaries as the target.
439        binary: Whether to use a binary segmentation target.
440        download: Whether to download the data if it is not present.
441        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
442
443    Returns:
444        The segmentation dataset.
445    """
446    if len(patch_shape) != 2:
447        raise ValueError(f"The Xenium patch shape must be two-dimensional, got {patch_shape}.")
448    if label_channel not in LABEL_CHANNELS:
449        raise ValueError(f"'{label_channel}' is not a valid label channel. Choose from {list(LABEL_CHANNELS)}.")
450    if raw_channel not in ("dapi", "stack"):
451        raise ValueError(f"'{raw_channel}' is not a valid raw channel. Choose from ['dapi', 'stack'].")
452
453    volume_paths = get_xenium_lung_treg_paths(path, sample, download, with_stains=raw_channel == "stack")
454
455    kwargs = util.update_kwargs(kwargs, "with_channels", raw_channel == "stack")
456    kwargs, _ = util.add_instance_label_transform(
457        kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary,
458    )
459    kwargs = util.ensure_transforms(ndim=2, **kwargs)
460
461    return torch_em.default_segmentation_dataset(
462        raw_paths=volume_paths,
463        raw_key=f"raw/{raw_channel}",
464        label_paths=volume_paths,
465        label_key=f"labels/{label_channel}",
466        patch_shape=patch_shape,
467        ndim=2,
468        **kwargs,
469    )

Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • patch_shape: The 2D patch shape to use for training.
  • sample: The sample or samples to use. Defaults to all of them.
  • label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
  • raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone, or 'stack' for all morphology channels.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • binary: Whether to use a binary segmentation target.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_xenium_lung_treg_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], sample: Union[str, Sequence[str], NoneType] = None, label_channel: Literal['nuclei', 'cells'] = 'nuclei', raw_channel: Literal['dapi', 'stack'] = 'dapi', offsets: Optional[List[List[int]]] = None, boundaries: bool = False, binary: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
472def get_xenium_lung_treg_loader(
473    path: Union[os.PathLike, str],
474    batch_size: int,
475    patch_shape: Tuple[int, int],
476    sample: Optional[Union[str, Sequence[str]]] = None,
477    label_channel: Literal["nuclei", "cells"] = "nuclei",
478    raw_channel: Literal["dapi", "stack"] = "dapi",
479    offsets: Optional[List[List[int]]] = None,
480    boundaries: bool = False,
481    binary: bool = False,
482    download: bool = False,
483    **kwargs,
484) -> DataLoader:
485    """Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation.
486
487    Args:
488        path: Filepath to a folder where the downloaded data will be saved.
489        batch_size: The batch size for training.
490        patch_shape: The 2D patch shape to use for training.
491        sample: The sample or samples to use. Defaults to all of them.
492        label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
493        raw_channel: The images to use as input. Either 'dapi' or 'stack'.
494        offsets: Offset values for affinity computation used as target.
495        boundaries: Whether to compute boundaries as the target.
496        binary: Whether to use a binary segmentation target.
497        download: Whether to download the data if it is not present.
498        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader.
499
500    Returns:
501        The DataLoader.
502    """
503    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
504    dataset = get_xenium_lung_treg_dataset(
505        path=path,
506        patch_shape=patch_shape,
507        sample=sample,
508        label_channel=label_channel,
509        raw_channel=raw_channel,
510        offsets=offsets,
511        boundaries=boundaries,
512        binary=binary,
513        download=download,
514        **ds_kwargs,
515    )
516    return torch_em.get_data_loader(dataset, batch_size=batch_size, **loader_kwargs)

Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • batch_size: The batch size for training.
  • patch_shape: The 2D patch shape to use for training.
  • sample: The sample or samples to use. Defaults to all of them.
  • label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
  • raw_channel: The images to use as input. Either 'dapi' or 'stack'.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • binary: Whether to use a binary segmentation target.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or the PyTorch DataLoader.
Returns:

The DataLoader.