torch_em.data.datasets.electron_microscopy.openorganelle_lipid_droplet

This dataset provides lipid droplet segmentation masks for representative crops from four mouse liver volumes on the public "OpenOrganelle" janelia-cosem-datasets S3 bucket (the same bucket used by cellmap.py, janelia_nucleus.py, and openorganelle_nucleus.py).

Three of the four datasets (jrc_mus-liver-4/5/6) are already covered by openorganelle_nucleus.py, but only for their nuclear-membrane (nuc_mem) label; their separate lipid droplet label, at labels/inference/segmentations/ld/{level}, is a distinct annotation not provided by that module. jrc_mus-liver-7 is not covered by any existing torch-em loader.

The lipid droplet label is a binary mask (only values 0/255 seen) that exactly matches the raw FIB-SEM data's multiscale pyramid at every level checked, so no resampling or axis correction is needed. Lipid droplets in these fatty-liver volumes can be very large (macrovesicular steatosis), so jrc_mus-liver-5 and jrc_mus-liver-6 use a coarser matched pyramid level (s3, 64nm) for their fixed crop, chosen to show a real droplet boundary rather than a single, deep-interior, fully-saturated block; jrc_mus-liver-4 and jrc_mus-liver-7 use the finest level (s0, 8nm).

The full matched-resolution raw+label pairs are far too large to download whole (hundreds of GB to over 1 TB); this module downloads one fixed, manually verified crop per dataset, matching the crop that was visually reviewed in napari before this loader was added. Data is CC0-1.0, per the bucket's established convention, except for jrc_mus-liver-7, whose license could not be independently confirmed (no dedicated landing page was found for this specific dataset).

Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this data.

  1"""This dataset provides lipid droplet segmentation masks for representative crops from four
  2mouse liver volumes on the public "OpenOrganelle" `janelia-cosem-datasets` S3 bucket (the same
  3bucket used by `cellmap.py`, `janelia_nucleus.py`, and `openorganelle_nucleus.py`).
  4
  5Three of the four datasets (`jrc_mus-liver-4/5/6`) are already covered by
  6`openorganelle_nucleus.py`, but only for their nuclear-membrane (`nuc_mem`) label; their separate
  7lipid droplet label, at `labels/inference/segmentations/ld/{level}`, is a distinct annotation not
  8provided by that module. `jrc_mus-liver-7` is not covered by any existing torch-em loader.
  9
 10The lipid droplet label is a binary mask (only values 0/255 seen) that exactly matches the raw
 11FIB-SEM data's multiscale pyramid at every level checked, so no resampling or axis correction is
 12needed. Lipid droplets in these fatty-liver volumes can be very large (macrovesicular steatosis),
 13so `jrc_mus-liver-5` and `jrc_mus-liver-6` use a coarser matched pyramid level (`s3`, 64nm) for
 14their fixed crop, chosen to show a real droplet boundary rather than a single, deep-interior,
 15fully-saturated block; `jrc_mus-liver-4` and `jrc_mus-liver-7` use the finest level (`s0`, 8nm).
 16
 17The full matched-resolution raw+label pairs are far too large to download whole (hundreds of GB
 18to over 1 TB); this module downloads one fixed, manually verified crop per dataset, matching the
 19crop that was visually reviewed in napari before this loader was added. Data is CC0-1.0, per the
 20bucket's established convention, except for `jrc_mus-liver-7`, whose license could not be
 21independently confirmed (no dedicated landing page was found for this specific dataset).
 22
 23Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this
 24data.
 25"""
 26
 27import os
 28from typing import List, Optional, Tuple, Union
 29
 30from torch.utils.data import DataLoader, Dataset
 31
 32import torch_em
 33
 34from .. import util
 35
 36
 37BUCKET_URL = "https://janelia-cosem-datasets.s3.amazonaws.com/"
 38LABEL_SUBPATH = "labels/inference/segmentations/ld"
 39
 40# dataset name -> recon folder, EM array name, ld pyramid level to use, and the bounding box
 41# (in that level's voxel coordinates) of the fixed, visually reviewed crop.
 42DATASETS = {
 43    "jrc_mus-liver-4": {
 44        "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s0",
 45        "bounding_box": ((704, 832), (1984, 2112), (4672, 4800)),
 46    },
 47    "jrc_mus-liver-5": {
 48        "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s3",
 49        "bounding_box": ((448, 640), (352, 544), (0, 128)),
 50    },
 51    "jrc_mus-liver-6": {
 52        "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s3",
 53        "bounding_box": ((288, 480), (0, 96), (736, 928)),
 54    },
 55    "jrc_mus-liver-7": {
 56        "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s0",
 57        "bounding_box": ((960, 1088), (5056, 5184), (8640, 8768)),
 58    },
 59}
 60
 61# `jrc_mus-liver-7` has no dedicated landing page confirming its license; the others follow the
 62# bucket's established CC0-1.0 convention (already verified for their existing torch-em entries).
 63UNVERIFIED_LICENSE = ("jrc_mus-liver-7",)
 64
 65
 66def _open_remote_zarr(s3_path):
 67    import zarr
 68    from zarr.storage import FsspecStore
 69
 70    # `FsspecStore` performs ranged partial reads, so a bounded crop only pulls the chunks it
 71    # actually needs, unlike `fsspec.get_mapper` which can fetch whole shard/chunk files.
 72    store = FsspecStore.from_url(s3_path, storage_options={"anon": True}, read_only=True)
 73    return zarr.open(store, mode="r")
 74
 75
 76def _get_json(url):
 77    import requests
 78
 79    resp = requests.get(url, timeout=60)
 80    resp.raise_for_status()
 81    return resp.json()
 82
 83
 84def _multiscale_levels(group_url):
 85    attrs = _get_json(f"{group_url}/.zattrs")
 86    return {
 87        d["path"]: tuple(d["coordinateTransformations"][0]["scale"])
 88        for d in attrs["multiscales"][0]["datasets"]
 89    }
 90
 91
 92def _find_matching_em_level(dataset_name, recon, em_name, label_scale):
 93    em_levels = _multiscale_levels(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}")
 94    for level, scale in em_levels.items():
 95        if scale == label_scale:
 96            return level
 97    raise RuntimeError(f"No EM pyramid level of '{dataset_name}' matches the ld label scale {label_scale}.")
 98
 99
100def get_openorganelle_lipid_droplet_data(
101    path: Union[os.PathLike, str], dataset_name: str, download: bool = False
102) -> str:
103    """Download the reviewed lipid droplet crop for one OpenOrganelle dataset.
104
105    Args:
106        path: Filepath to a folder where the cached zarr store will be saved.
107        dataset_name: The name of the dataset, one of the keys in `DATASETS`.
108        download: Whether to download the data if it is not present.
109
110    Returns:
111        The filepath to the cached zarr store.
112    """
113    import zarr
114    from zarr.codecs import BloscCodec
115
116    if dataset_name not in DATASETS:
117        raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.")
118
119    info = DATASETS[dataset_name]
120    recon, em_name, label_level = info["recon"], info["em_name"], info["label_level"]
121    bounding_box = info["bounding_box"]
122
123    os.makedirs(str(path), exist_ok=True)
124    zarr_path = os.path.join(str(path), f"{dataset_name}.zarr")
125
126    root = zarr.open_group(zarr_path, mode="a")
127    if "raw" in root and "labels" in root:
128        return zarr_path
129
130    if not download:
131        raise RuntimeError(f"No cached crop found at '{zarr_path}'. Set download=True to stream it from S3.")
132
133    print(f"Streaming a lipid droplet crop for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...")
134    label_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}/{label_level}"
135    label_arr = _open_remote_zarr(label_url)
136    label_slices = tuple(slice(*bb) for bb in bounding_box)
137    label_data = label_arr[label_slices]
138
139    label_scale = _multiscale_levels(
140        f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}"
141    )[label_level]
142    em_level = _find_matching_em_level(dataset_name, recon, em_name, label_scale)
143    em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}"
144    em_arr = _open_remote_zarr(em_url)
145    raw_data = em_arr[label_slices]
146
147    common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape))
148    raw_data = raw_data[tuple(slice(0, s) for s in common_shape)]
149    label_data = label_data[tuple(slice(0, s) for s in common_shape)]
150
151    assert raw_data.shape == label_data.shape, (
152        f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}"
153    )
154
155    def _make_array(name, data, shuffle):
156        arr = root.create_array(
157            name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype,
158            compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle),
159        )
160        arr[:] = data
161
162    root.attrs["dataset"] = dataset_name
163    root.attrs["recon"] = recon
164    root.attrs["label_target"] = "lipid_droplet"
165    root.attrs["label_source"] = label_url
166    root.attrs["em_source"] = em_url
167    root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE
168
169    _make_array("raw", raw_data, shuffle="shuffle")
170    _make_array("labels", label_data, shuffle="bitshuffle")
171
172    print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).")
173    return zarr_path
174
175
176def get_openorganelle_lipid_droplet_paths(
177    path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False,
178) -> List[str]:
179    """Get paths to cached OpenOrganelle lipid droplet zarr stores.
180
181    Args:
182        path: Filepath to a folder where the cached zarr stores will be saved.
183        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
184        download: Whether to download the data if it is not present.
185
186    Returns:
187        List of filepaths to the cached zarr stores.
188    """
189    if dataset_names is None:
190        dataset_names = list(DATASETS.keys())
191    return [get_openorganelle_lipid_droplet_data(path, name, download) for name in dataset_names]
192
193
194def get_openorganelle_lipid_droplet_dataset(
195    path: Union[os.PathLike, str],
196    patch_shape: Tuple[int, int, int],
197    dataset_names: Optional[List[str]] = None,
198    download: bool = False,
199    **kwargs,
200) -> Dataset:
201    """Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation.
202
203    Args:
204        path: Filepath to a folder where the cached zarr stores will be saved.
205        patch_shape: The patch shape (z, y, x) to use for training.
206        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
207        download: Whether to download the data if it is not present.
208        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
209
210    Returns:
211        The segmentation dataset.
212    """
213    assert len(patch_shape) == 3
214
215    paths = get_openorganelle_lipid_droplet_paths(path, dataset_names, download)
216
217    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
218
219    return torch_em.default_segmentation_dataset(
220        raw_paths=paths,
221        raw_key="raw",
222        label_paths=paths,
223        label_key="labels",
224        patch_shape=patch_shape,
225        **kwargs,
226    )
227
228
229def get_openorganelle_lipid_droplet_loader(
230    path: Union[os.PathLike, str],
231    patch_shape: Tuple[int, int, int],
232    batch_size: int,
233    dataset_names: Optional[List[str]] = None,
234    download: bool = False,
235    **kwargs,
236) -> DataLoader:
237    """Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset.
238
239    Args:
240        path: Filepath to a folder where the cached zarr stores will be saved.
241        patch_shape: The patch shape (z, y, x) to use for training.
242        batch_size: The batch size for training.
243        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
244        download: Whether to download the data if it is not present.
245        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
246            or for the PyTorch DataLoader.
247
248    Returns:
249        The DataLoader.
250    """
251    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
252    dataset = get_openorganelle_lipid_droplet_dataset(
253        path, patch_shape, dataset_names=dataset_names, download=download, **ds_kwargs
254    )
255    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
BUCKET_URL = 'https://janelia-cosem-datasets.s3.amazonaws.com/'
LABEL_SUBPATH = 'labels/inference/segmentations/ld'
DATASETS = {'jrc_mus-liver-4': {'recon': 'recon-1', 'em_name': 'fibsem-uint8', 'label_level': 's0', 'bounding_box': ((704, 832), (1984, 2112), (4672, 4800))}, 'jrc_mus-liver-5': {'recon': 'recon-1', 'em_name': 'fibsem-uint8', 'label_level': 's3', 'bounding_box': ((448, 640), (352, 544), (0, 128))}, 'jrc_mus-liver-6': {'recon': 'recon-1', 'em_name': 'fibsem-uint8', 'label_level': 's3', 'bounding_box': ((288, 480), (0, 96), (736, 928))}, 'jrc_mus-liver-7': {'recon': 'recon-1', 'em_name': 'fibsem-uint8', 'label_level': 's0', 'bounding_box': ((960, 1088), (5056, 5184), (8640, 8768))}}
UNVERIFIED_LICENSE = ('jrc_mus-liver-7',)
def get_openorganelle_lipid_droplet_data( path: Union[os.PathLike, str], dataset_name: str, download: bool = False) -> str:
101def get_openorganelle_lipid_droplet_data(
102    path: Union[os.PathLike, str], dataset_name: str, download: bool = False
103) -> str:
104    """Download the reviewed lipid droplet crop for one OpenOrganelle dataset.
105
106    Args:
107        path: Filepath to a folder where the cached zarr store will be saved.
108        dataset_name: The name of the dataset, one of the keys in `DATASETS`.
109        download: Whether to download the data if it is not present.
110
111    Returns:
112        The filepath to the cached zarr store.
113    """
114    import zarr
115    from zarr.codecs import BloscCodec
116
117    if dataset_name not in DATASETS:
118        raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.")
119
120    info = DATASETS[dataset_name]
121    recon, em_name, label_level = info["recon"], info["em_name"], info["label_level"]
122    bounding_box = info["bounding_box"]
123
124    os.makedirs(str(path), exist_ok=True)
125    zarr_path = os.path.join(str(path), f"{dataset_name}.zarr")
126
127    root = zarr.open_group(zarr_path, mode="a")
128    if "raw" in root and "labels" in root:
129        return zarr_path
130
131    if not download:
132        raise RuntimeError(f"No cached crop found at '{zarr_path}'. Set download=True to stream it from S3.")
133
134    print(f"Streaming a lipid droplet crop for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...")
135    label_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}/{label_level}"
136    label_arr = _open_remote_zarr(label_url)
137    label_slices = tuple(slice(*bb) for bb in bounding_box)
138    label_data = label_arr[label_slices]
139
140    label_scale = _multiscale_levels(
141        f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}"
142    )[label_level]
143    em_level = _find_matching_em_level(dataset_name, recon, em_name, label_scale)
144    em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}"
145    em_arr = _open_remote_zarr(em_url)
146    raw_data = em_arr[label_slices]
147
148    common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape))
149    raw_data = raw_data[tuple(slice(0, s) for s in common_shape)]
150    label_data = label_data[tuple(slice(0, s) for s in common_shape)]
151
152    assert raw_data.shape == label_data.shape, (
153        f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}"
154    )
155
156    def _make_array(name, data, shuffle):
157        arr = root.create_array(
158            name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype,
159            compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle),
160        )
161        arr[:] = data
162
163    root.attrs["dataset"] = dataset_name
164    root.attrs["recon"] = recon
165    root.attrs["label_target"] = "lipid_droplet"
166    root.attrs["label_source"] = label_url
167    root.attrs["em_source"] = em_url
168    root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE
169
170    _make_array("raw", raw_data, shuffle="shuffle")
171    _make_array("labels", label_data, shuffle="bitshuffle")
172
173    print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).")
174    return zarr_path

Download the reviewed lipid droplet crop for one OpenOrganelle dataset.

Arguments:
  • path: Filepath to a folder where the cached zarr store will be saved.
  • dataset_name: The name of the dataset, one of the keys in DATASETS.
  • download: Whether to download the data if it is not present.
Returns:

The filepath to the cached zarr store.

def get_openorganelle_lipid_droplet_paths( path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False) -> List[str]:
177def get_openorganelle_lipid_droplet_paths(
178    path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False,
179) -> List[str]:
180    """Get paths to cached OpenOrganelle lipid droplet zarr stores.
181
182    Args:
183        path: Filepath to a folder where the cached zarr stores will be saved.
184        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
185        download: Whether to download the data if it is not present.
186
187    Returns:
188        List of filepaths to the cached zarr stores.
189    """
190    if dataset_names is None:
191        dataset_names = list(DATASETS.keys())
192    return [get_openorganelle_lipid_droplet_data(path, name, download) for name in dataset_names]

Get paths to cached OpenOrganelle lipid droplet zarr stores.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • dataset_names: The names of the datasets to use. Defaults to all datasets in DATASETS.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths to the cached zarr stores.

def get_openorganelle_lipid_droplet_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], dataset_names: Optional[List[str]] = None, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
195def get_openorganelle_lipid_droplet_dataset(
196    path: Union[os.PathLike, str],
197    patch_shape: Tuple[int, int, int],
198    dataset_names: Optional[List[str]] = None,
199    download: bool = False,
200    **kwargs,
201) -> Dataset:
202    """Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation.
203
204    Args:
205        path: Filepath to a folder where the cached zarr stores will be saved.
206        patch_shape: The patch shape (z, y, x) to use for training.
207        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
208        download: Whether to download the data if it is not present.
209        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
210
211    Returns:
212        The segmentation dataset.
213    """
214    assert len(patch_shape) == 3
215
216    paths = get_openorganelle_lipid_droplet_paths(path, dataset_names, download)
217
218    kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True)
219
220    return torch_em.default_segmentation_dataset(
221        raw_paths=paths,
222        raw_key="raw",
223        label_paths=paths,
224        label_key="labels",
225        patch_shape=patch_shape,
226        **kwargs,
227    )

Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • dataset_names: The names of the datasets to use. Defaults to all datasets in DATASETS.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_openorganelle_lipid_droplet_loader( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], batch_size: int, dataset_names: Optional[List[str]] = None, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
230def get_openorganelle_lipid_droplet_loader(
231    path: Union[os.PathLike, str],
232    patch_shape: Tuple[int, int, int],
233    batch_size: int,
234    dataset_names: Optional[List[str]] = None,
235    download: bool = False,
236    **kwargs,
237) -> DataLoader:
238    """Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset.
239
240    Args:
241        path: Filepath to a folder where the cached zarr stores will be saved.
242        patch_shape: The patch shape (z, y, x) to use for training.
243        batch_size: The batch size for training.
244        dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`.
245        download: Whether to download the data if it is not present.
246        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`
247            or for the PyTorch DataLoader.
248
249    Returns:
250        The DataLoader.
251    """
252    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
253    dataset = get_openorganelle_lipid_droplet_dataset(
254        path, patch_shape, dataset_names=dataset_names, download=download, **ds_kwargs
255    )
256    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset.

Arguments:
  • path: Filepath to a folder where the cached zarr stores will be saved.
  • patch_shape: The patch shape (z, y, x) to use for training.
  • batch_size: The batch size for training.
  • dataset_names: The names of the datasets to use. Defaults to all datasets in DATASETS.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.