torch_em.data.datasets.electron_microscopy.pombe_nucleus_mito

This dataset provides nucleus, mitochondrion, and lipid droplet segmentation masks for cryo-ET tomograms of Schizosaccharomyces pombe cells.

The data is hosted on the CryoET Data Portal, part of the DeePiCt benchmark family - the same paper family as the actin-only deepict.py loader, but different datasets within it (deepict.py hardcodes dataset 10002, the actin dataset). This module covers:

  • Dataset 10001 (https://cryoetdataportal.czscience.com/datasets/10001, EMPIAR-10988), run "TS_0006": nucleus and mitochondrion.
  • Dataset 10000 (https://cryoetdataportal.czscience.com/datasets/10000), run "TS_045": nucleus and mitochondrion. Runs "TS_028" and "TS_041": lipid droplet.

Both datasets have many more runs and several more annotated structures (cytoplasm, vesicle, endoplasmic reticulum, nuclear envelope, Golgi apparatus, membrane); only the runs and targets listed above are provided here, since that is what was verified and visually reviewed before this loader was added.

All masks are real, expert-verified ground truth (groundTruthStatus=true, methodType=hybrid per the portal's own metadata), not automated predictions. Data is CC0-1.0, per the CryoET Data Portal's portal-wide terms of use. The corresponding author for dataset 10000, Julia Mahamid, has a public email (julia.mahamid@embl.de); no public email could be found for Judith B. Zaugg or for dataset 10001's corresponding authors despite a real search.

  1"""This dataset provides nucleus, mitochondrion, and lipid droplet segmentation masks for cryo-ET
  2tomograms of Schizosaccharomyces pombe cells.
  3
  4The data is hosted on the CryoET Data Portal, part of the DeePiCt benchmark family - the same
  5paper family as the actin-only `deepict.py` loader, but different datasets within it (`deepict.py`
  6hardcodes dataset 10002, the actin dataset). This module covers:
  7
  8- Dataset 10001 (https://cryoetdataportal.czscience.com/datasets/10001, EMPIAR-10988), run
  9  "TS_0006": nucleus and mitochondrion.
 10- Dataset 10000 (https://cryoetdataportal.czscience.com/datasets/10000), run "TS_045": nucleus
 11  and mitochondrion. Runs "TS_028" and "TS_041": lipid droplet.
 12
 13Both datasets have many more runs and several more annotated structures (cytoplasm, vesicle,
 14endoplasmic reticulum, nuclear envelope, Golgi apparatus, membrane); only the runs and targets
 15listed above are provided here, since that is what was verified and visually reviewed before
 16this loader was added.
 17
 18All masks are real, expert-verified ground truth (`groundTruthStatus=true`, `methodType=hybrid`
 19per the portal's own metadata), not automated predictions. Data is CC0-1.0, per the CryoET Data
 20Portal's portal-wide terms of use. The corresponding author for dataset 10000, Julia Mahamid, has
 21a public email (julia.mahamid@embl.de); no public email could be found for Judith B. Zaugg or for
 22dataset 10001's corresponding authors despite a real search.
 23"""
 24
 25import os
 26from typing import List, Tuple, Union
 27
 28from torch.utils.data import DataLoader, Dataset
 29
 30import torch_em
 31
 32from .. import util
 33
 34
 35BUCKET = "cryoet-data-portal-public"
 36
 37# (dataset_id, run) -> {target: annotation folder}. A run's zarr group caches "raw" plus one
 38# array per target listed here, so a run contributing multiple targets (e.g. TS_045's nucleus and
 39# mitochondrion) is only downloaded once regardless of which target is requested.
 40RUN_LABELS = {
 41    (10001, "TS_0006"): {"nucleus": "108", "mitochondrion": "103"},
 42    (10000, "TS_045"): {"nucleus": "108", "mitochondrion": "103"},
 43    (10000, "TS_028"): {"lipid_droplet": "110"},
 44    (10000, "TS_041"): {"lipid_droplet": "110"},
 45}
 46TARGETS = ("nucleus", "mitochondrion", "lipid_droplet")
 47
 48
 49def _open_remote_zarr(s3_path):
 50    import zarr
 51    from zarr.storage import FsspecStore
 52
 53    store = FsspecStore.from_url(s3_path, storage_options={"anon": True}, read_only=True)
 54    return zarr.open(store, mode="r")
 55
 56
 57def _run_prefix(dataset_id, run):
 58    return f"s3://{BUCKET}/{dataset_id}/{run}/Reconstructions/VoxelSpacing13.480"
 59
 60
 61def get_pombe_nucleus_mito_data(
 62    path: Union[os.PathLike, str], dataset_id: int, run: str, download: bool = False
 63) -> str:
 64    """Download one S. pombe cryo-ET run's tomogram and its available organelle masks.
 65
 66    Args:
 67        path: Filepath to a folder where the cached zarr store will be saved.
 68        dataset_id: The CryoET Data Portal dataset id, one of the keys in `RUN_LABELS`.
 69        run: The run name, matching `dataset_id` in `RUN_LABELS`.
 70        download: Whether to download the data if it is not present.
 71
 72    Returns:
 73        The filepath to the cached zarr store.
 74    """
 75    import zarr
 76    from zarr.codecs import BloscCodec
 77
 78    key = (dataset_id, run)
 79    if key not in RUN_LABELS:
 80        raise ValueError(f"'{key}' is not a valid (dataset_id, run). Choose from {sorted(RUN_LABELS.keys())}.")
 81    labels = RUN_LABELS[key]
 82
 83    os.makedirs(str(path), exist_ok=True)
 84    zarr_path = os.path.join(str(path), f"{dataset_id}_{run}.zarr")
 85
 86    root = zarr.open_group(zarr_path, mode="a")
 87    if "raw" in root and all(name in root for name in labels):
 88        return zarr_path
 89
 90    if not download:
 91        raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.")
 92
 93    print(f"Streaming the S. pombe {run} tomogram and masks from the CryoET Data Portal ...")
 94    prefix = _run_prefix(dataset_id, run)
 95    raw = _open_remote_zarr(f"{prefix}/Tomograms/100/{run}.zarr/0")[:]
 96
 97    def _make_array(name, data, shuffle):
 98        arr = root.create_array(
 99            name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype,
100            compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle),
101        )
102        arr[:] = data
103
104    if "raw" not in root:
105        _make_array("raw", raw, shuffle="shuffle")
106
107    for name, folder in labels.items():
108        if name in root:
109            continue
110        label_data = _open_remote_zarr(f"{prefix}/Annotations/{folder}/{name}-1.0_segmentationmask.zarr/0")[:]
111        assert label_data.shape == raw.shape, (
112            f"Shape mismatch for {key}, target '{name}': raw {raw.shape} vs label {label_data.shape}"
113        )
114        _make_array(name, label_data, shuffle="bitshuffle")
115
116    root.attrs["dataset_id"] = dataset_id
117    root.attrs["run"] = run
118    root.attrs["ground_truth"] = True
119
120    print(f"Cached the S. pombe {run} data to '{zarr_path}' (shape {raw.shape}).")
121    return zarr_path
122
123
124def get_pombe_nucleus_mito_paths(
125    path: Union[os.PathLike, str], target: str = "nucleus", download: bool = False,
126) -> List[str]:
127    """Get paths to cached S. pombe zarr stores that provide the requested target.
128
129    Args:
130        path: Filepath to a folder where the cached zarr stores will be saved.
131        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
132        download: Whether to download the data if it is not present.
133
134    Returns:
135        List of filepaths to the cached zarr stores.
136    """
137    if target not in TARGETS:
138        raise ValueError(f"'target' must be one of {TARGETS}, got '{target}'.")
139    runs = [key for key, labels in RUN_LABELS.items() if target in labels]
140    return [get_pombe_nucleus_mito_data(path, dataset_id, run, download) for dataset_id, run in runs]
141
142
143def get_pombe_nucleus_mito_dataset(
144    path: Union[os.PathLike, str],
145    patch_shape: Tuple[int, int, int],
146    target: str = "nucleus",
147    download: bool = False,
148    **kwargs,
149) -> Dataset:
150    """Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation.
151
152    Args:
153        path: Filepath to a folder where the cached zarr stores will be saved.
154        patch_shape: The patch shape (z, y, x) to use for training.
155        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
156        download: Whether to download the data if it is not present.
157        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
158
159    Returns:
160        The segmentation dataset.
161    """
162    assert len(patch_shape) == 3
163
164    paths = get_pombe_nucleus_mito_paths(path, target, download)
165
166    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
167
168    return torch_em.default_segmentation_dataset(
169        raw_paths=paths,
170        raw_key="raw",
171        label_paths=paths,
172        label_key=target,
173        patch_shape=patch_shape,
174        **kwargs,
175    )
176
177
178def get_pombe_nucleus_mito_loader(
179    path: Union[os.PathLike, str],
180    patch_shape: Tuple[int, int, int],
181    batch_size: int,
182    target: str = "nucleus",
183    download: bool = False,
184    **kwargs,
185) -> DataLoader:
186    """Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset.
187
188    Args:
189        path: Filepath to a folder where the cached zarr stores will be saved.
190        patch_shape: The patch shape (z, y, x) to use for training.
191        batch_size: The batch size for training.
192        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
193        download: Whether to download the data if it is not present.
194        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
195            or for the PyTorch DataLoader.
196
197    Returns:
198        The DataLoader.
199    """
200    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
201    dataset = get_pombe_nucleus_mito_dataset(path, patch_shape, target=target, download=download, **ds_kwargs)
202    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
BUCKET = 'cryoet-data-portal-public'
RUN_LABELS = {(10001, 'TS_0006'): {'nucleus': '108', 'mitochondrion': '103'}, (10000, 'TS_045'): {'nucleus': '108', 'mitochondrion': '103'}, (10000, 'TS_028'): {'lipid_droplet': '110'}, (10000, 'TS_041'): {'lipid_droplet': '110'}}
TARGETS = ('nucleus', 'mitochondrion', 'lipid_droplet')
def get_pombe_nucleus_mito_data( path: Union[os.PathLike, str], dataset_id: int, run: str, download: bool = False) -> str:
 62def get_pombe_nucleus_mito_data(
 63    path: Union[os.PathLike, str], dataset_id: int, run: str, download: bool = False
 64) -> str:
 65    """Download one S. pombe cryo-ET run's tomogram and its available organelle masks.
 66
 67    Args:
 68        path: Filepath to a folder where the cached zarr store will be saved.
 69        dataset_id: The CryoET Data Portal dataset id, one of the keys in `RUN_LABELS`.
 70        run: The run name, matching `dataset_id` in `RUN_LABELS`.
 71        download: Whether to download the data if it is not present.
 72
 73    Returns:
 74        The filepath to the cached zarr store.
 75    """
 76    import zarr
 77    from zarr.codecs import BloscCodec
 78
 79    key = (dataset_id, run)
 80    if key not in RUN_LABELS:
 81        raise ValueError(f"'{key}' is not a valid (dataset_id, run). Choose from {sorted(RUN_LABELS.keys())}.")
 82    labels = RUN_LABELS[key]
 83
 84    os.makedirs(str(path), exist_ok=True)
 85    zarr_path = os.path.join(str(path), f"{dataset_id}_{run}.zarr")
 86
 87    root = zarr.open_group(zarr_path, mode="a")
 88    if "raw" in root and all(name in root for name in labels):
 89        return zarr_path
 90
 91    if not download:
 92        raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.")
 93
 94    print(f"Streaming the S. pombe {run} tomogram and masks from the CryoET Data Portal ...")
 95    prefix = _run_prefix(dataset_id, run)
 96    raw = _open_remote_zarr(f"{prefix}/Tomograms/100/{run}.zarr/0")[:]
 97
 98    def _make_array(name, data, shuffle):
 99        arr = root.create_array(
100            name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype,
101            compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle),
102        )
103        arr[:] = data
104
105    if "raw" not in root:
106        _make_array("raw", raw, shuffle="shuffle")
107
108    for name, folder in labels.items():
109        if name in root:
110            continue
111        label_data = _open_remote_zarr(f"{prefix}/Annotations/{folder}/{name}-1.0_segmentationmask.zarr/0")[:]
112        assert label_data.shape == raw.shape, (
113            f"Shape mismatch for {key}, target '{name}': raw {raw.shape} vs label {label_data.shape}"
114        )
115        _make_array(name, label_data, shuffle="bitshuffle")
116
117    root.attrs["dataset_id"] = dataset_id
118    root.attrs["run"] = run
119    root.attrs["ground_truth"] = True
120
121    print(f"Cached the S. pombe {run} data to '{zarr_path}' (shape {raw.shape}).")
122    return zarr_path

Download one S. pombe cryo-ET run's tomogram and its available organelle masks.

Arguments:
  • path: Filepath to a folder where the cached zarr store will be saved.
  • dataset_id: The CryoET Data Portal dataset id, one of the keys in RUN_LABELS.
  • run: The run name, matching dataset_id in RUN_LABELS.
  • download: Whether to download the data if it is not present.
Returns:

The filepath to the cached zarr store.

def get_pombe_nucleus_mito_paths( path: Union[os.PathLike, str], target: str = 'nucleus', download: bool = False) -> List[str]:
125def get_pombe_nucleus_mito_paths(
126    path: Union[os.PathLike, str], target: str = "nucleus", download: bool = False,
127) -> List[str]:
128    """Get paths to cached S. pombe zarr stores that provide the requested target.
129
130    Args:
131        path: Filepath to a folder where the cached zarr stores will be saved.
132        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
133        download: Whether to download the data if it is not present.
134
135    Returns:
136        List of filepaths to the cached zarr stores.
137    """
138    if target not in TARGETS:
139        raise ValueError(f"'target' must be one of {TARGETS}, got '{target}'.")
140    runs = [key for key, labels in RUN_LABELS.items() if target in labels]
141    return [get_pombe_nucleus_mito_data(path, dataset_id, run, download) for dataset_id, run in runs]

Get paths to cached S. pombe zarr stores that provide the requested target.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths to the cached zarr stores.

def get_pombe_nucleus_mito_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], target: str = 'nucleus', download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
144def get_pombe_nucleus_mito_dataset(
145    path: Union[os.PathLike, str],
146    patch_shape: Tuple[int, int, int],
147    target: str = "nucleus",
148    download: bool = False,
149    **kwargs,
150) -> Dataset:
151    """Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation.
152
153    Args:
154        path: Filepath to a folder where the cached zarr stores will be saved.
155        patch_shape: The patch shape (z, y, x) to use for training.
156        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
157        download: Whether to download the data if it is not present.
158        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
159
160    Returns:
161        The segmentation dataset.
162    """
163    assert len(patch_shape) == 3
164
165    paths = get_pombe_nucleus_mito_paths(path, target, download)
166
167    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
168
169    return torch_em.default_segmentation_dataset(
170        raw_paths=paths,
171        raw_key="raw",
172        label_paths=paths,
173        label_key=target,
174        patch_shape=patch_shape,
175        **kwargs,
176    )

Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_pombe_nucleus_mito_loader( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], batch_size: int, target: str = 'nucleus', download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
179def get_pombe_nucleus_mito_loader(
180    path: Union[os.PathLike, str],
181    patch_shape: Tuple[int, int, int],
182    batch_size: int,
183    target: str = "nucleus",
184    download: bool = False,
185    **kwargs,
186) -> DataLoader:
187    """Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset.
188
189    Args:
190        path: Filepath to a folder where the cached zarr stores will be saved.
191        patch_shape: The patch shape (z, y, x) to use for training.
192        batch_size: The batch size for training.
193        target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
194        download: Whether to download the data if it is not present.
195        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
196            or for the PyTorch DataLoader.
197
198    Returns:
199        The DataLoader.
200    """
201    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
202    dataset = get_pombe_nucleus_mito_dataset(path, patch_shape, target=target, download=download, **ds_kwargs)
203    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • batch_size: The batch size for training.
  • target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.