torch_em.data.datasets.electron_microscopy.cmito

Cmito dataset for dense mitochondria segmentation in whole-animal C. elegans EM volumes.

The dataset provides dense mitochondria instance masks for five developmental stages of C. elegans, built on public whole-animal EM connectome volumes:

  • l1, l2, l3, adult: from Witvliet et al. 2021 (https://doi.org/10.1038/s41586-021-03778-8), CC-BY-4.0, deposited at https://zenodo.org/records/5637219 and https://bossdb.org/project/witvliet2020.
  • dauer: from Yim et al. 2024 (https://doi.org/10.1038/s41467-024-45943-3), whose article text is CC-BY-4.0. The mitochondria mask bucket for this stage lives in a separate GCS project (gnd-dauer1) than the other four stages, and no repository-level data license could be confirmed for it independently of the article license. Treat "dauer" as license-uncertain until a repository landing page or downloadable asset states its data license explicitly.

Raw EM is streamed from BossDB (S3) at a mip level chosen per stage so that it spatially matches the coarser, separately hosted mitochondria mask (on Google Cloud Storage, at native resolution). Both are cached locally as zarr v3 stores in (z, y, x) axis order.

This dataset is from the publication https://doi.org/10.1101/2024.07.19.604219. Please cite it if you use this dataset in your research.

Corresponding authors: J. Alexander Bae (jabae@snu.ac.kr), Junho Lee (elegans@snu.ac.kr). Original tool repository: https://github.com/jabae/Cmito. Requires cloud-volume: pip install cloud-volume.

  1"""Cmito dataset for dense mitochondria segmentation in whole-animal C. elegans EM volumes.
  2
  3The dataset provides dense mitochondria instance masks for five developmental stages of
  4C. elegans, built on public whole-animal EM connectome volumes:
  5
  6- l1, l2, l3, adult: from Witvliet et al. 2021 (https://doi.org/10.1038/s41586-021-03778-8),
  7  CC-BY-4.0, deposited at https://zenodo.org/records/5637219 and https://bossdb.org/project/witvliet2020.
  8- dauer: from Yim et al. 2024 (https://doi.org/10.1038/s41467-024-45943-3), whose article text
  9  is CC-BY-4.0. The mitochondria mask bucket for this stage lives in a separate GCS project
 10  (`gnd-dauer1`) than the other four stages, and no repository-level data license could be
 11  confirmed for it independently of the article license. Treat "dauer" as license-uncertain
 12  until a repository landing page or downloadable asset states its data license explicitly.
 13
 14Raw EM is streamed from BossDB (S3) at a mip level chosen per stage so that it spatially matches
 15the coarser, separately hosted mitochondria mask (on Google Cloud Storage, at native resolution).
 16Both are cached locally as zarr v3 stores in (z, y, x) axis order.
 17
 18This dataset is from the publication https://doi.org/10.1101/2024.07.19.604219.
 19Please cite it if you use this dataset in your research.
 20
 21Corresponding authors: J. Alexander Bae (jabae@snu.ac.kr), Junho Lee (elegans@snu.ac.kr).
 22Original tool repository: https://github.com/jabae/Cmito.
 23Requires cloud-volume: pip install cloud-volume.
 24"""
 25
 26import hashlib
 27import os
 28from concurrent.futures import ThreadPoolExecutor, as_completed
 29from typing import List, Literal, Optional, Sequence, Tuple, Union
 30
 31import numpy as np
 32from tqdm import tqdm
 33from torch.utils.data import Dataset, DataLoader
 34
 35import torch_em
 36from .. import util
 37
 38
 39CMITO_STAGES = {
 40    "l1": {
 41        "raw_url": "precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_2/em",
 42        "mask_url": "precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset2/mito_seg_v3",
 43        # Raw mip3 (0.64 nm x 8 = 5.12 nm xy) spatially matches the mask's native 5.12/5.12/50 nm.
 44        "raw_mip": 3,
 45        # Full extent in nm, from the mip0 shape (26624, 22016, 368) x resolution (0.64, 0.64, 50).
 46        "bbox_nm": (0, 17039, 0, 14090, 0, 18400),
 47    },
 48    "l2": {
 49        "raw_url": "precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_5/em",
 50        "mask_url": "precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset5/mito_seg_v4",
 51        # Raw mip2 (2 nm x 4 = 8 nm xy) spatially matches the mask's native 8/8/30 nm.
 52        "raw_mip": 2,
 53        "bbox_nm": (0, 27648, 0, 22528, 0, 25920),
 54    },
 55    "l3": {
 56        "raw_url": "precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_6/em",
 57        "mask_url": "precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset6/mito_seg_v4",
 58        # Raw mip3 (8x downsample) spatially matches the mask, despite the mask's reported
 59        # 8 nm nominal resolution not being an exact 8x multiple of the raw 0.768 nm mip0
 60        # pixel size; shapes (4544 x 4288) confirm the 8x factor is the correct match.
 61        "raw_mip": 3,
 62        "bbox_nm": (0, 27918, 0, 26345, 0, 21600),
 63    },
 64    "adult": {
 65        "raw_url": "precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_8/em",
 66        "mask_url": "precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset8/mito_seg_v3",
 67        # Raw mip3 (2 nm x 8 = 16 nm xy) spatially matches the mask's native 16/16/30 nm.
 68        "raw_mip": 3,
 69        "bbox_nm": (0, 79872, 0, 44032, 0, 21120),
 70    },
 71    "dauer": {
 72        "raw_url": "precomputed://https://bossdb-open-data.s3.amazonaws.com/yim_choe_bae2023/dauer1_364/em/em",
 73        # NOTE: different GCS bucket/project than the other four stages.
 74        "mask_url": "precomputed://https://storage.googleapis.com/gnd-dauer1/dauer1_364/mito_seg_v4",
 75        # Raw mip3 (1 nm x 8 = 8 nm xy) spatially matches the mask's native 8/8/50 nm.
 76        "raw_mip": 3,
 77        "bbox_nm": (0, 9800, 0, 9792, 0, 18200),
 78    },
 79}
 80
 81CMITO_CHUNK_SHAPE = (64, 128, 128)
 82CMITO_SHARD_SHAPE = (128, 512, 512)
 83
 84
 85def _cmito_bbox_to_str(bbox):
 86    return hashlib.md5("_".join(str(v) for v in bbox).encode()).hexdigest()[:12]
 87
 88
 89def _cmito_create_array(root, name, shape, dtype, is_label):
 90    from zarr.codecs import BloscCodec
 91    shuffle = "bitshuffle" if (np.issubdtype(dtype, np.integer) and is_label) else "shuffle"
 92    return root.create_array(
 93        name,
 94        shape=shape,
 95        chunks=CMITO_CHUNK_SHAPE,
 96        shards=CMITO_SHARD_SHAPE,
 97        dtype=dtype,
 98        compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle),
 99    )
100
101
102def _cmito_bbox_voxels(cv, x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm):
103    scale = np.array(cv.resolution)
104    x0 = int(np.floor(x_min_nm / scale[0]))
105    x1 = int(np.ceil(x_max_nm / scale[0]))
106    y0 = int(np.floor(y_min_nm / scale[1]))
107    y1 = int(np.ceil(y_max_nm / scale[1]))
108    z0 = int(np.floor(z_min_nm / scale[2]))
109    z1 = int(np.ceil(z_max_nm / scale[2]))
110    return x0, x1, y0, y1, z0, z1, (z1 - z0, y1 - y0, x1 - x0)
111
112
113def _cmito_download_to_zarr(cv, ds, x0g, y0g, z0g, name):
114    shape = ds.shape  # (z, y, x)
115    sz, sy, sx = CMITO_SHARD_SHAPE
116
117    tasks = []
118    for z0_ in range(0, shape[0], sz):
119        for y0_ in range(0, shape[1], sy):
120            for x0_ in range(0, shape[2], sx):
121                z1_ = min(z0_ + sz, shape[0])
122                y1_ = min(y0_ + sy, shape[1])
123                x1_ = min(x0_ + sx, shape[2])
124                tasks.append((
125                    (z0_, z1_), (y0_, y1_), (x0_, x1_),
126                    (x0g + x0_, x0g + x1_, y0g + y0_, y0g + y1_, z0g + z0_, z0g + z1_),
127                ))
128
129    target_dtype = np.dtype(ds.dtype)
130
131    def worker(item):
132        (z0_, z1_), (y0_, y1_), (x0_, x1_), (gx0, gx1, gy0, gy1, gz0, gz1) = item
133        block = np.asarray(cv[gx0:gx1, gy0:gy1, gz0:gz1])
134        if block.ndim == 4:
135            block = block[..., 0]
136        ds[z0_:z1_, y0_:y1_, x0_:x1_] = block.transpose(2, 1, 0).astype(target_dtype)
137
138    with ThreadPoolExecutor(max_workers=8) as ex:
139        futures = [ex.submit(worker, t) for t in tasks]
140        for fut in tqdm(as_completed(futures), total=len(futures), desc=f"Downloading '{name}'", smoothing=0.05):
141            fut.result()
142
143
144def get_cmito_data(
145    path: Union[os.PathLike, str],
146    stage: Literal["l1", "l2", "l3", "adult", "dauer"],
147    bounding_box: Optional[Tuple[float, ...]] = None,
148    download: bool = False,
149) -> str:
150    """Stream and cache one Cmito developmental-stage volume as a zarr v3 store.
151
152    The zarr store contains:
153      - raw: EM grayscale (uint8, z/y/x), at the mip level that spatially matches the mask.
154      - labels: mitochondria instance segmentation (uint16, z/y/x), at native mask resolution.
155
156    Args:
157        path: Filepath to a folder where the cached zarr store will be saved.
158        stage: The developmental stage to use. One of 'l1', 'l2', 'l3', 'adult', 'dauer'.
159            The 'dauer' stage's mask bucket has no independently confirmed data license
160            (see the module docstring); use it with that caveat in mind.
161        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max).
162            Defaults to the full volume extent for the chosen stage.
163        download: Whether to stream and cache the data if not present.
164
165    Returns:
166        Filepath to the cached zarr store.
167    """
168    import zarr
169
170    if stage not in CMITO_STAGES:
171        raise ValueError(f"Invalid stage: '{stage}'. Choose from {list(CMITO_STAGES.keys())}.")
172
173    stage_info = CMITO_STAGES[stage]
174    os.makedirs(str(path), exist_ok=True)
175    bbox = bounding_box if bounding_box is not None else stage_info["bbox_nm"]
176    bbox_hash = _cmito_bbox_to_str(bbox)
177    zarr_path = os.path.join(str(path), f"{stage}_{bbox_hash}.zarr")
178
179    def _complete(zp):
180        return os.path.isdir(os.path.join(zp, "raw")) and os.path.isdir(os.path.join(zp, "labels"))
181
182    if _complete(zarr_path):
183        return zarr_path
184    if not download:
185        raise RuntimeError(
186            f"No cached data at '{zarr_path}'. Set download=True to stream from BossDB and GCS."
187        )
188
189    try:
190        from cloudvolume import CloudVolume
191    except ImportError:
192        raise ImportError("The 'cloud-volume' package is required: pip install cloud-volume")
193
194    x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm = bbox
195    raw_mip = stage_info["raw_mip"]
196    print(f"Streaming Cmito stage='{stage}' at raw_mip={raw_mip} ...")
197
198    raw_cv = CloudVolume(
199        stage_info["raw_url"], use_https=True, mip=raw_mip, progress=False, fill_missing=True,
200    )
201    mask_cv = CloudVolume(stage_info["mask_url"], mip=0, progress=False, fill_missing=True)
202
203    rx0, rx1, ry0, ry1, rz0, rz1, raw_shape = _cmito_bbox_voxels(
204        raw_cv, x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm
205    )
206    mx0, mx1, my0, my1, mz0, mz1, mask_shape = _cmito_bbox_voxels(
207        mask_cv, x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm
208    )
209    shape = tuple(min(r, m) for r, m in zip(raw_shape, mask_shape))
210
211    root = zarr.open_group(zarr_path, mode="a")
212    root.attrs["stage"] = stage
213    root.attrs["bounding_box_nm"] = list(bbox)
214    root.attrs["raw_mip"] = raw_mip
215
216    if "raw" not in root:
217        ds_raw = _cmito_create_array(root, "raw", shape, np.dtype("uint8"), is_label=False)
218        _cmito_download_to_zarr(raw_cv, ds_raw, rx0, ry0, rz0, name="raw")
219
220    if "labels" not in root:
221        ds_lbl = _cmito_create_array(root, "labels", shape, np.dtype("uint16"), is_label=True)
222        _cmito_download_to_zarr(mask_cv, ds_lbl, mx0, my0, mz0, name="labels")
223
224    print(f"Cached to {zarr_path} (shape {shape})")
225    return zarr_path
226
227
228def get_cmito_paths(
229    path: Union[os.PathLike, str],
230    stages: Optional[Sequence[str]] = None,
231    bounding_box: Optional[Tuple[float, ...]] = None,
232    download: bool = False,
233) -> List[str]:
234    """Get paths to cached Cmito zarr stores.
235
236    Args:
237        path: Filepath to a folder where the cached zarr stores will be saved.
238        stages: Developmental stages to load. Defaults to all five stages.
239        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
240            every requested stage. Defaults to each stage's own full volume extent.
241        download: Whether to stream and cache the data if not present.
242
243    Returns:
244        Filepaths to the cached zarr stores.
245    """
246    stages_ = list(stages) if stages is not None else list(CMITO_STAGES.keys())
247    return [get_cmito_data(path, stage, bounding_box, download) for stage in stages_]
248
249
250def get_cmito_dataset(
251    path: Union[os.PathLike, str],
252    patch_shape: Tuple[int, int, int],
253    stage: Union[str, Sequence[str]] = "l1",
254    bounding_box: Optional[Tuple[float, ...]] = None,
255    download: bool = False,
256    offsets: Optional[List[List[int]]] = None,
257    boundaries: bool = False,
258    **kwargs,
259) -> Dataset:
260    """Get the Cmito dataset for mitochondria instance segmentation in whole-animal C. elegans EM.
261
262    Args:
263        path: Filepath to a folder where the cached zarr stores will be saved.
264        patch_shape: The patch shape (z, y, x) to use for training.
265        stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult',
266            'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license
267            (see the module docstring); use it with that caveat in mind.
268        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
269            every requested stage. Defaults to each stage's own full volume extent.
270        download: Whether to stream and cache data if not already present.
271        offsets: Offset values for affinity computation used as target.
272        boundaries: Whether to compute boundaries as the target.
273        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
274
275    Returns:
276        The segmentation dataset.
277    """
278    assert len(patch_shape) == 3
279    stages = [stage] if isinstance(stage, str) else list(stage)
280    paths = get_cmito_paths(path, stages, bounding_box, download)
281
282    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
283    kwargs, _ = util.add_instance_label_transform(
284        kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets
285    )
286
287    return torch_em.default_segmentation_dataset(
288        raw_paths=paths,
289        raw_key="raw",
290        label_paths=paths,
291        label_key="labels",
292        patch_shape=patch_shape,
293        **kwargs,
294    )
295
296
297def get_cmito_loader(
298    path: Union[os.PathLike, str],
299    batch_size: int,
300    patch_shape: Tuple[int, int, int],
301    stage: Union[str, Sequence[str]] = "l1",
302    bounding_box: Optional[Tuple[float, ...]] = None,
303    download: bool = False,
304    offsets: Optional[List[List[int]]] = None,
305    boundaries: bool = False,
306    **kwargs,
307) -> DataLoader:
308    """Get the DataLoader for mitochondria instance segmentation in whole-animal C. elegans EM.
309
310    Args:
311        path: Filepath to a folder where the cached zarr stores will be saved.
312        batch_size: The batch size for training.
313        patch_shape: The patch shape (z, y, x) to use for training.
314        stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult',
315            'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license
316            (see the module docstring); use it with that caveat in mind.
317        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
318            every requested stage. Defaults to each stage's own full volume extent.
319        download: Whether to stream and cache data if not already present.
320        offsets: Offset values for affinity computation used as target.
321        boundaries: Whether to compute boundaries as the target.
322        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
323            or for the PyTorch DataLoader.
324
325    Returns:
326        The DataLoader.
327    """
328    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
329    ds = get_cmito_dataset(
330        path=path,
331        patch_shape=patch_shape,
332        stage=stage,
333        bounding_box=bounding_box,
334        download=download,
335        offsets=offsets,
336        boundaries=boundaries,
337        **ds_kwargs,
338    )
339    return torch_em.get_data_loader(ds, batch_size=batch_size, **loader_kwargs)
CMITO_STAGES = {'l1': {'raw_url': 'precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_2/em', 'mask_url': 'precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset2/mito_seg_v3', 'raw_mip': 3, 'bbox_nm': (0, 17039, 0, 14090, 0, 18400)}, 'l2': {'raw_url': 'precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_5/em', 'mask_url': 'precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset5/mito_seg_v4', 'raw_mip': 2, 'bbox_nm': (0, 27648, 0, 22528, 0, 25920)}, 'l3': {'raw_url': 'precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_6/em', 'mask_url': 'precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset6/mito_seg_v4', 'raw_mip': 3, 'bbox_nm': (0, 27918, 0, 26345, 0, 21600)}, 'adult': {'raw_url': 'precomputed://https://bossdb-open-data.s3.amazonaws.com/witvliet2020/Dataset_8/em', 'mask_url': 'precomputed://https://storage.googleapis.com/gnd-neuroglancer/witvliet/dataset8/mito_seg_v3', 'raw_mip': 3, 'bbox_nm': (0, 79872, 0, 44032, 0, 21120)}, 'dauer': {'raw_url': 'precomputed://https://bossdb-open-data.s3.amazonaws.com/yim_choe_bae2023/dauer1_364/em/em', 'mask_url': 'precomputed://https://storage.googleapis.com/gnd-dauer1/dauer1_364/mito_seg_v4', 'raw_mip': 3, 'bbox_nm': (0, 9800, 0, 9792, 0, 18200)}}
CMITO_CHUNK_SHAPE = (64, 128, 128)
CMITO_SHARD_SHAPE = (128, 512, 512)
def get_cmito_data( path: Union[os.PathLike, str], stage: Literal['l1', 'l2', 'l3', 'adult', 'dauer'], bounding_box: Optional[Tuple[float, ...]] = None, download: bool = False) -> str:
145def get_cmito_data(
146    path: Union[os.PathLike, str],
147    stage: Literal["l1", "l2", "l3", "adult", "dauer"],
148    bounding_box: Optional[Tuple[float, ...]] = None,
149    download: bool = False,
150) -> str:
151    """Stream and cache one Cmito developmental-stage volume as a zarr v3 store.
152
153    The zarr store contains:
154      - raw: EM grayscale (uint8, z/y/x), at the mip level that spatially matches the mask.
155      - labels: mitochondria instance segmentation (uint16, z/y/x), at native mask resolution.
156
157    Args:
158        path: Filepath to a folder where the cached zarr store will be saved.
159        stage: The developmental stage to use. One of 'l1', 'l2', 'l3', 'adult', 'dauer'.
160            The 'dauer' stage's mask bucket has no independently confirmed data license
161            (see the module docstring); use it with that caveat in mind.
162        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max).
163            Defaults to the full volume extent for the chosen stage.
164        download: Whether to stream and cache the data if not present.
165
166    Returns:
167        Filepath to the cached zarr store.
168    """
169    import zarr
170
171    if stage not in CMITO_STAGES:
172        raise ValueError(f"Invalid stage: '{stage}'. Choose from {list(CMITO_STAGES.keys())}.")
173
174    stage_info = CMITO_STAGES[stage]
175    os.makedirs(str(path), exist_ok=True)
176    bbox = bounding_box if bounding_box is not None else stage_info["bbox_nm"]
177    bbox_hash = _cmito_bbox_to_str(bbox)
178    zarr_path = os.path.join(str(path), f"{stage}_{bbox_hash}.zarr")
179
180    def _complete(zp):
181        return os.path.isdir(os.path.join(zp, "raw")) and os.path.isdir(os.path.join(zp, "labels"))
182
183    if _complete(zarr_path):
184        return zarr_path
185    if not download:
186        raise RuntimeError(
187            f"No cached data at '{zarr_path}'. Set download=True to stream from BossDB and GCS."
188        )
189
190    try:
191        from cloudvolume import CloudVolume
192    except ImportError:
193        raise ImportError("The 'cloud-volume' package is required: pip install cloud-volume")
194
195    x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm = bbox
196    raw_mip = stage_info["raw_mip"]
197    print(f"Streaming Cmito stage='{stage}' at raw_mip={raw_mip} ...")
198
199    raw_cv = CloudVolume(
200        stage_info["raw_url"], use_https=True, mip=raw_mip, progress=False, fill_missing=True,
201    )
202    mask_cv = CloudVolume(stage_info["mask_url"], mip=0, progress=False, fill_missing=True)
203
204    rx0, rx1, ry0, ry1, rz0, rz1, raw_shape = _cmito_bbox_voxels(
205        raw_cv, x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm
206    )
207    mx0, mx1, my0, my1, mz0, mz1, mask_shape = _cmito_bbox_voxels(
208        mask_cv, x_min_nm, x_max_nm, y_min_nm, y_max_nm, z_min_nm, z_max_nm
209    )
210    shape = tuple(min(r, m) for r, m in zip(raw_shape, mask_shape))
211
212    root = zarr.open_group(zarr_path, mode="a")
213    root.attrs["stage"] = stage
214    root.attrs["bounding_box_nm"] = list(bbox)
215    root.attrs["raw_mip"] = raw_mip
216
217    if "raw" not in root:
218        ds_raw = _cmito_create_array(root, "raw", shape, np.dtype("uint8"), is_label=False)
219        _cmito_download_to_zarr(raw_cv, ds_raw, rx0, ry0, rz0, name="raw")
220
221    if "labels" not in root:
222        ds_lbl = _cmito_create_array(root, "labels", shape, np.dtype("uint16"), is_label=True)
223        _cmito_download_to_zarr(mask_cv, ds_lbl, mx0, my0, mz0, name="labels")
224
225    print(f"Cached to {zarr_path} (shape {shape})")
226    return zarr_path

Stream and cache one Cmito developmental-stage volume as a zarr v3 store.

The zarr store contains:
  • raw: EM grayscale (uint8, z/y/x), at the mip level that spatially matches the mask.
  • labels: mitochondria instance segmentation (uint16, z/y/x), at native mask resolution.
Arguments:
  • path: Filepath to a folder where the cached zarr store will be saved.
  • stage: The developmental stage to use. One of 'l1', 'l2', 'l3', 'adult', 'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license (see the module docstring); use it with that caveat in mind.
  • bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max). Defaults to the full volume extent for the chosen stage.
  • download: Whether to stream and cache the data if not present.
Returns:

Filepath to the cached zarr store.

def get_cmito_paths( path: Union[os.PathLike, str], stages: Optional[Sequence[str]] = None, bounding_box: Optional[Tuple[float, ...]] = None, download: bool = False) -> List[str]:
229def get_cmito_paths(
230    path: Union[os.PathLike, str],
231    stages: Optional[Sequence[str]] = None,
232    bounding_box: Optional[Tuple[float, ...]] = None,
233    download: bool = False,
234) -> List[str]:
235    """Get paths to cached Cmito zarr stores.
236
237    Args:
238        path: Filepath to a folder where the cached zarr stores will be saved.
239        stages: Developmental stages to load. Defaults to all five stages.
240        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
241            every requested stage. Defaults to each stage's own full volume extent.
242        download: Whether to stream and cache the data if not present.
243
244    Returns:
245        Filepaths to the cached zarr stores.
246    """
247    stages_ = list(stages) if stages is not None else list(CMITO_STAGES.keys())
248    return [get_cmito_data(path, stage, bounding_box, download) for stage in stages_]

Get paths to cached Cmito zarr stores.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • stages: Developmental stages to load. Defaults to all five stages.
  • bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to every requested stage. Defaults to each stage's own full volume extent.
  • download: Whether to stream and cache the data if not present.
Returns:

Filepaths to the cached zarr stores.

def get_cmito_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], stage: Union[str, Sequence[str]] = 'l1', bounding_box: Optional[Tuple[float, ...]] = None, download: bool = False, offsets: Optional[List[List[int]]] = None, boundaries: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
251def get_cmito_dataset(
252    path: Union[os.PathLike, str],
253    patch_shape: Tuple[int, int, int],
254    stage: Union[str, Sequence[str]] = "l1",
255    bounding_box: Optional[Tuple[float, ...]] = None,
256    download: bool = False,
257    offsets: Optional[List[List[int]]] = None,
258    boundaries: bool = False,
259    **kwargs,
260) -> Dataset:
261    """Get the Cmito dataset for mitochondria instance segmentation in whole-animal C. elegans EM.
262
263    Args:
264        path: Filepath to a folder where the cached zarr stores will be saved.
265        patch_shape: The patch shape (z, y, x) to use for training.
266        stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult',
267            'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license
268            (see the module docstring); use it with that caveat in mind.
269        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
270            every requested stage. Defaults to each stage's own full volume extent.
271        download: Whether to stream and cache data if not already present.
272        offsets: Offset values for affinity computation used as target.
273        boundaries: Whether to compute boundaries as the target.
274        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
275
276    Returns:
277        The segmentation dataset.
278    """
279    assert len(patch_shape) == 3
280    stages = [stage] if isinstance(stage, str) else list(stage)
281    paths = get_cmito_paths(path, stages, bounding_box, download)
282
283    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
284    kwargs, _ = util.add_instance_label_transform(
285        kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets
286    )
287
288    return torch_em.default_segmentation_dataset(
289        raw_paths=paths,
290        raw_key="raw",
291        label_paths=paths,
292        label_key="labels",
293        patch_shape=patch_shape,
294        **kwargs,
295    )

Get the Cmito dataset for mitochondria instance segmentation in whole-animal C. elegans EM.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult', 'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license (see the module docstring); use it with that caveat in mind.
  • bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to every requested stage. Defaults to each stage's own full volume extent.
  • download: Whether to stream and cache data if not already present.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_cmito_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int, int], stage: Union[str, Sequence[str]] = 'l1', bounding_box: Optional[Tuple[float, ...]] = None, download: bool = False, offsets: Optional[List[List[int]]] = None, boundaries: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
298def get_cmito_loader(
299    path: Union[os.PathLike, str],
300    batch_size: int,
301    patch_shape: Tuple[int, int, int],
302    stage: Union[str, Sequence[str]] = "l1",
303    bounding_box: Optional[Tuple[float, ...]] = None,
304    download: bool = False,
305    offsets: Optional[List[List[int]]] = None,
306    boundaries: bool = False,
307    **kwargs,
308) -> DataLoader:
309    """Get the DataLoader for mitochondria instance segmentation in whole-animal C. elegans EM.
310
311    Args:
312        path: Filepath to a folder where the cached zarr stores will be saved.
313        batch_size: The batch size for training.
314        patch_shape: The patch shape (z, y, x) to use for training.
315        stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult',
316            'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license
317            (see the module docstring); use it with that caveat in mind.
318        bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to
319            every requested stage. Defaults to each stage's own full volume extent.
320        download: Whether to stream and cache data if not already present.
321        offsets: Offset values for affinity computation used as target.
322        boundaries: Whether to compute boundaries as the target.
323        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
324            or for the PyTorch DataLoader.
325
326    Returns:
327        The DataLoader.
328    """
329    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
330    ds = get_cmito_dataset(
331        path=path,
332        patch_shape=patch_shape,
333        stage=stage,
334        bounding_box=bounding_box,
335        download=download,
336        offsets=offsets,
337        boundaries=boundaries,
338        **ds_kwargs,
339    )
340    return torch_em.get_data_loader(ds, batch_size=batch_size, **loader_kwargs)

Get the DataLoader for mitochondria instance segmentation in whole-animal C. elegans EM.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • stage: The developmental stage(s) to use, one or several of 'l1', 'l2', 'l3', 'adult', 'dauer'. The 'dauer' stage's mask bucket has no independently confirmed data license (see the module docstring); use it with that caveat in mind.
  • bounding_box: Region in nm as (x_min, x_max, y_min, y_max, z_min, z_max), applied to every requested stage. Defaults to each stage's own full volume extent.
  • download: Whether to stream and cache data if not already present.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.