torch_em.data.datasets.electron_microscopy.pombe_nucleus_mito
This dataset provides nucleus, mitochondrion, and lipid droplet segmentation masks for cryo-ET tomograms of Schizosaccharomyces pombe cells.
The data is hosted on the CryoET Data Portal, part of the DeePiCt benchmark family - the same
paper family as the actin-only deepict.py loader, but different datasets within it (deepict.py
hardcodes dataset 10002, the actin dataset). This module covers:
- Dataset 10001 (https://cryoetdataportal.czscience.com/datasets/10001, EMPIAR-10988), run "TS_0006": nucleus and mitochondrion.
- Dataset 10000 (https://cryoetdataportal.czscience.com/datasets/10000), run "TS_045": nucleus and mitochondrion. Runs "TS_028" and "TS_041": lipid droplet.
Both datasets have many more runs and several more annotated structures (cytoplasm, vesicle, endoplasmic reticulum, nuclear envelope, Golgi apparatus, membrane); only the runs and targets listed above are provided here, since that is what was verified and visually reviewed before this loader was added.
All masks are real, expert-verified ground truth (groundTruthStatus=true, methodType=hybrid
per the portal's own metadata), not automated predictions. Data is CC0-1.0, per the CryoET Data
Portal's portal-wide terms of use. The corresponding author for dataset 10000, Julia Mahamid, has
a public email (julia.mahamid@embl.de); no public email could be found for Judith B. Zaugg or for
dataset 10001's corresponding authors despite a real search.
1"""This dataset provides nucleus, mitochondrion, and lipid droplet segmentation masks for cryo-ET 2tomograms of Schizosaccharomyces pombe cells. 3 4The data is hosted on the CryoET Data Portal, part of the DeePiCt benchmark family - the same 5paper family as the actin-only `deepict.py` loader, but different datasets within it (`deepict.py` 6hardcodes dataset 10002, the actin dataset). This module covers: 7 8- Dataset 10001 (https://cryoetdataportal.czscience.com/datasets/10001, EMPIAR-10988), run 9 "TS_0006": nucleus and mitochondrion. 10- Dataset 10000 (https://cryoetdataportal.czscience.com/datasets/10000), run "TS_045": nucleus 11 and mitochondrion. Runs "TS_028" and "TS_041": lipid droplet. 12 13Both datasets have many more runs and several more annotated structures (cytoplasm, vesicle, 14endoplasmic reticulum, nuclear envelope, Golgi apparatus, membrane); only the runs and targets 15listed above are provided here, since that is what was verified and visually reviewed before 16this loader was added. 17 18All masks are real, expert-verified ground truth (`groundTruthStatus=true`, `methodType=hybrid` 19per the portal's own metadata), not automated predictions. Data is CC0-1.0, per the CryoET Data 20Portal's portal-wide terms of use. The corresponding author for dataset 10000, Julia Mahamid, has 21a public email (julia.mahamid@embl.de); no public email could be found for Judith B. Zaugg or for 22dataset 10001's corresponding authors despite a real search. 23""" 24 25import os 26from typing import List, Tuple, Union 27 28from torch.utils.data import DataLoader, Dataset 29 30import torch_em 31 32from .. import util 33 34 35BUCKET = "cryoet-data-portal-public" 36 37# (dataset_id, run) -> {target: annotation folder}. A run's zarr group caches "raw" plus one 38# array per target listed here, so a run contributing multiple targets (e.g. TS_045's nucleus and 39# mitochondrion) is only downloaded once regardless of which target is requested. 40RUN_LABELS = { 41 (10001, "TS_0006"): {"nucleus": "108", "mitochondrion": "103"}, 42 (10000, "TS_045"): {"nucleus": "108", "mitochondrion": "103"}, 43 (10000, "TS_028"): {"lipid_droplet": "110"}, 44 (10000, "TS_041"): {"lipid_droplet": "110"}, 45} 46TARGETS = ("nucleus", "mitochondrion", "lipid_droplet") 47 48 49def _open_remote_zarr(s3_path): 50 import zarr 51 from zarr.storage import FsspecStore 52 53 store = FsspecStore.from_url(s3_path, storage_options={"anon": True}, read_only=True) 54 return zarr.open(store, mode="r") 55 56 57def _run_prefix(dataset_id, run): 58 return f"s3://{BUCKET}/{dataset_id}/{run}/Reconstructions/VoxelSpacing13.480" 59 60 61def get_pombe_nucleus_mito_data( 62 path: Union[os.PathLike, str], dataset_id: int, run: str, download: bool = False 63) -> str: 64 """Download one S. pombe cryo-ET run's tomogram and its available organelle masks. 65 66 Args: 67 path: Filepath to a folder where the cached zarr store will be saved. 68 dataset_id: The CryoET Data Portal dataset id, one of the keys in `RUN_LABELS`. 69 run: The run name, matching `dataset_id` in `RUN_LABELS`. 70 download: Whether to download the data if it is not present. 71 72 Returns: 73 The filepath to the cached zarr store. 74 """ 75 import zarr 76 from zarr.codecs import BloscCodec 77 78 key = (dataset_id, run) 79 if key not in RUN_LABELS: 80 raise ValueError(f"'{key}' is not a valid (dataset_id, run). Choose from {sorted(RUN_LABELS.keys())}.") 81 labels = RUN_LABELS[key] 82 83 os.makedirs(str(path), exist_ok=True) 84 zarr_path = os.path.join(str(path), f"{dataset_id}_{run}.zarr") 85 86 root = zarr.open_group(zarr_path, mode="a") 87 if "raw" in root and all(name in root for name in labels): 88 return zarr_path 89 90 if not download: 91 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 92 93 print(f"Streaming the S. pombe {run} tomogram and masks from the CryoET Data Portal ...") 94 prefix = _run_prefix(dataset_id, run) 95 raw = _open_remote_zarr(f"{prefix}/Tomograms/100/{run}.zarr/0")[:] 96 97 def _make_array(name, data, shuffle): 98 arr = root.create_array( 99 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 100 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 101 ) 102 arr[:] = data 103 104 if "raw" not in root: 105 _make_array("raw", raw, shuffle="shuffle") 106 107 for name, folder in labels.items(): 108 if name in root: 109 continue 110 label_data = _open_remote_zarr(f"{prefix}/Annotations/{folder}/{name}-1.0_segmentationmask.zarr/0")[:] 111 assert label_data.shape == raw.shape, ( 112 f"Shape mismatch for {key}, target '{name}': raw {raw.shape} vs label {label_data.shape}" 113 ) 114 _make_array(name, label_data, shuffle="bitshuffle") 115 116 root.attrs["dataset_id"] = dataset_id 117 root.attrs["run"] = run 118 root.attrs["ground_truth"] = True 119 120 print(f"Cached the S. pombe {run} data to '{zarr_path}' (shape {raw.shape}).") 121 return zarr_path 122 123 124def get_pombe_nucleus_mito_paths( 125 path: Union[os.PathLike, str], target: str = "nucleus", download: bool = False, 126) -> List[str]: 127 """Get paths to cached S. pombe zarr stores that provide the requested target. 128 129 Args: 130 path: Filepath to a folder where the cached zarr stores will be saved. 131 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 132 download: Whether to download the data if it is not present. 133 134 Returns: 135 List of filepaths to the cached zarr stores. 136 """ 137 if target not in TARGETS: 138 raise ValueError(f"'target' must be one of {TARGETS}, got '{target}'.") 139 runs = [key for key, labels in RUN_LABELS.items() if target in labels] 140 return [get_pombe_nucleus_mito_data(path, dataset_id, run, download) for dataset_id, run in runs] 141 142 143def get_pombe_nucleus_mito_dataset( 144 path: Union[os.PathLike, str], 145 patch_shape: Tuple[int, int, int], 146 target: str = "nucleus", 147 download: bool = False, 148 **kwargs, 149) -> Dataset: 150 """Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation. 151 152 Args: 153 path: Filepath to a folder where the cached zarr stores will be saved. 154 patch_shape: The patch shape (z, y, x) to use for training. 155 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 156 download: Whether to download the data if it is not present. 157 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 158 159 Returns: 160 The segmentation dataset. 161 """ 162 assert len(patch_shape) == 3 163 164 paths = get_pombe_nucleus_mito_paths(path, target, download) 165 166 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 167 168 return torch_em.default_segmentation_dataset( 169 raw_paths=paths, 170 raw_key="raw", 171 label_paths=paths, 172 label_key=target, 173 patch_shape=patch_shape, 174 **kwargs, 175 ) 176 177 178def get_pombe_nucleus_mito_loader( 179 path: Union[os.PathLike, str], 180 patch_shape: Tuple[int, int, int], 181 batch_size: int, 182 target: str = "nucleus", 183 download: bool = False, 184 **kwargs, 185) -> DataLoader: 186 """Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset. 187 188 Args: 189 path: Filepath to a folder where the cached zarr stores will be saved. 190 patch_shape: The patch shape (z, y, x) to use for training. 191 batch_size: The batch size for training. 192 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 193 download: Whether to download the data if it is not present. 194 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 195 or for the PyTorch DataLoader. 196 197 Returns: 198 The DataLoader. 199 """ 200 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 201 dataset = get_pombe_nucleus_mito_dataset(path, patch_shape, target=target, download=download, **ds_kwargs) 202 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
62def get_pombe_nucleus_mito_data( 63 path: Union[os.PathLike, str], dataset_id: int, run: str, download: bool = False 64) -> str: 65 """Download one S. pombe cryo-ET run's tomogram and its available organelle masks. 66 67 Args: 68 path: Filepath to a folder where the cached zarr store will be saved. 69 dataset_id: The CryoET Data Portal dataset id, one of the keys in `RUN_LABELS`. 70 run: The run name, matching `dataset_id` in `RUN_LABELS`. 71 download: Whether to download the data if it is not present. 72 73 Returns: 74 The filepath to the cached zarr store. 75 """ 76 import zarr 77 from zarr.codecs import BloscCodec 78 79 key = (dataset_id, run) 80 if key not in RUN_LABELS: 81 raise ValueError(f"'{key}' is not a valid (dataset_id, run). Choose from {sorted(RUN_LABELS.keys())}.") 82 labels = RUN_LABELS[key] 83 84 os.makedirs(str(path), exist_ok=True) 85 zarr_path = os.path.join(str(path), f"{dataset_id}_{run}.zarr") 86 87 root = zarr.open_group(zarr_path, mode="a") 88 if "raw" in root and all(name in root for name in labels): 89 return zarr_path 90 91 if not download: 92 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 93 94 print(f"Streaming the S. pombe {run} tomogram and masks from the CryoET Data Portal ...") 95 prefix = _run_prefix(dataset_id, run) 96 raw = _open_remote_zarr(f"{prefix}/Tomograms/100/{run}.zarr/0")[:] 97 98 def _make_array(name, data, shuffle): 99 arr = root.create_array( 100 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 101 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 102 ) 103 arr[:] = data 104 105 if "raw" not in root: 106 _make_array("raw", raw, shuffle="shuffle") 107 108 for name, folder in labels.items(): 109 if name in root: 110 continue 111 label_data = _open_remote_zarr(f"{prefix}/Annotations/{folder}/{name}-1.0_segmentationmask.zarr/0")[:] 112 assert label_data.shape == raw.shape, ( 113 f"Shape mismatch for {key}, target '{name}': raw {raw.shape} vs label {label_data.shape}" 114 ) 115 _make_array(name, label_data, shuffle="bitshuffle") 116 117 root.attrs["dataset_id"] = dataset_id 118 root.attrs["run"] = run 119 root.attrs["ground_truth"] = True 120 121 print(f"Cached the S. pombe {run} data to '{zarr_path}' (shape {raw.shape}).") 122 return zarr_path
Download one S. pombe cryo-ET run's tomogram and its available organelle masks.
Arguments:
- path: Filepath to a folder where the cached zarr store will be saved.
- dataset_id: The CryoET Data Portal dataset id, one of the keys in
RUN_LABELS. - run: The run name, matching
dataset_idinRUN_LABELS. - download: Whether to download the data if it is not present.
Returns:
The filepath to the cached zarr store.
125def get_pombe_nucleus_mito_paths( 126 path: Union[os.PathLike, str], target: str = "nucleus", download: bool = False, 127) -> List[str]: 128 """Get paths to cached S. pombe zarr stores that provide the requested target. 129 130 Args: 131 path: Filepath to a folder where the cached zarr stores will be saved. 132 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 133 download: Whether to download the data if it is not present. 134 135 Returns: 136 List of filepaths to the cached zarr stores. 137 """ 138 if target not in TARGETS: 139 raise ValueError(f"'target' must be one of {TARGETS}, got '{target}'.") 140 runs = [key for key, labels in RUN_LABELS.items() if target in labels] 141 return [get_pombe_nucleus_mito_data(path, dataset_id, run, download) for dataset_id, run in runs]
Get paths to cached S. pombe zarr stores that provide the requested target.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths to the cached zarr stores.
144def get_pombe_nucleus_mito_dataset( 145 path: Union[os.PathLike, str], 146 patch_shape: Tuple[int, int, int], 147 target: str = "nucleus", 148 download: bool = False, 149 **kwargs, 150) -> Dataset: 151 """Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation. 152 153 Args: 154 path: Filepath to a folder where the cached zarr stores will be saved. 155 patch_shape: The patch shape (z, y, x) to use for training. 156 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 157 download: Whether to download the data if it is not present. 158 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 159 160 Returns: 161 The segmentation dataset. 162 """ 163 assert len(patch_shape) == 3 164 165 paths = get_pombe_nucleus_mito_paths(path, target, download) 166 167 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 168 169 return torch_em.default_segmentation_dataset( 170 raw_paths=paths, 171 raw_key="raw", 172 label_paths=paths, 173 label_key=target, 174 patch_shape=patch_shape, 175 **kwargs, 176 )
Get the S. pombe dataset for nucleus, mitochondrion, or lipid droplet segmentation.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
179def get_pombe_nucleus_mito_loader( 180 path: Union[os.PathLike, str], 181 patch_shape: Tuple[int, int, int], 182 batch_size: int, 183 target: str = "nucleus", 184 download: bool = False, 185 **kwargs, 186) -> DataLoader: 187 """Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset. 188 189 Args: 190 path: Filepath to a folder where the cached zarr stores will be saved. 191 patch_shape: The patch shape (z, y, x) to use for training. 192 batch_size: The batch size for training. 193 target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'. 194 download: Whether to download the data if it is not present. 195 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 196 or for the PyTorch DataLoader. 197 198 Returns: 199 The DataLoader. 200 """ 201 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 202 dataset = get_pombe_nucleus_mito_dataset(path, patch_shape, target=target, download=download, **ds_kwargs) 203 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DataLoader for nucleus, mitochondrion, or lipid droplet segmentation in the S. pombe dataset.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- batch_size: The batch size for training.
- target: The segmentation target, one of 'nucleus', 'mitochondrion', or 'lipid_droplet'.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.