torch_em.data.datasets.light_microscopy.slice2

The SLICE-2 dataset contains annotations for 3D nucleus instance segmentation in light-sheet fluorescence microscopy volumes of Tribolium castaneum embryos.

SLICE-2 (Second Systematic Live Imaging Collection of Embryogenesis) is a collection of sixteen isotropic 3D live imaging datasets of gastrulation and early germband elongation of the red flour beetle, imaged with a histone-labeled (H2A/H2B) transgenic line. The nuclei of selected time points are segmented, so that this module provides 15 labeled volumes from 14 of the datasets (DATASETS lists the labeled time points of each). The raw data are the deconvolved z-stacks ('(D1)-ZStacks-Decon-ZS', about 600 x 1100 x 600 voxels, uint16), and the labels are nucleus instance ids on the same grid. The segmentation is stored by the authors in a different axis order (1000 x 600 x 600, flipped along one axis), which get_slice2_data converts. It was checked against the raw volumes, where the segmented voxels are 4 - 8 times brighter than the background.

The published segmentation contains objects with the reserved id 65535 (about 6% of the segmented voxels, many small fragments). These voxels are set to background and the remaining ids are made consecutive. The segmentation of DS0001, DS0003 and further time points is only distributed as colored 'CH(TM)' RGB stacks, which do not match the raw grid of the instance stacks. They are not used by this module.

NOTE: Every volume takes about 0.55 GB (raw) plus the labels to download. The raw data are stored in ~35 GB zip archives on Zenodo (one per dataset), which are never downloaded as a whole: the members are read with HTTP range requests and checked against the CRC32 of the archive. Use datasets to only download a subset of the datasets.

The data is located at https://doi.org/10.5281/zenodo.18015798 (segmentations) and in one Zenodo record per raw archive (see DATASETS), released under a CC-BY-4.0 license.

This dataset is from the publication Kraemer et al. (2026), see the Zenodo records for the details. Please cite it if you use this dataset for your research.

  1"""The SLICE-2 dataset contains annotations for 3D nucleus instance segmentation in light-sheet fluorescence
  2microscopy volumes of Tribolium castaneum embryos.
  3
  4SLICE-2 (Second Systematic Live Imaging Collection of Embryogenesis) is a collection of sixteen isotropic 3D
  5live imaging datasets of gastrulation and early germband elongation of the red flour beetle, imaged with a
  6histone-labeled (H2A/H2B) transgenic line. The nuclei of selected time points are segmented, so that this module
  7provides 15 labeled volumes from 14 of the datasets (`DATASETS` lists the labeled time points of each). The
  8raw data are the deconvolved z-stacks ('(D1)-ZStacks-Decon-ZS', about 600 x 1100 x 600 voxels, uint16), and the labels
  9are nucleus instance ids on the same grid. The segmentation is stored by the authors in a different axis order
 10(1000 x 600 x 600, flipped along one axis), which `get_slice2_data` converts. It was checked against the raw
 11volumes, where the segmented voxels are 4 - 8 times brighter than the background.
 12
 13The published segmentation contains objects with the reserved id 65535 (about 6% of the segmented voxels, many small
 14fragments). These voxels are set to background and the remaining ids are made consecutive. The segmentation of
 15DS0001, DS0003 and further time points is only distributed as colored 'CH(TM)' RGB stacks, which do not match the
 16raw grid of the instance stacks. They are not used by this module.
 17
 18NOTE: Every volume takes about 0.55 GB (raw) plus the labels to download. The raw data are stored in ~35 GB zip
 19archives on Zenodo (one per dataset), which are never downloaded as a whole: the members are read with HTTP range
 20requests and checked against the CRC32 of the archive. Use `datasets` to only download a subset of the datasets.
 21
 22The data is located at https://doi.org/10.5281/zenodo.18015798 (segmentations) and in one Zenodo record per raw
 23archive (see `DATASETS`), released under a CC-BY-4.0 license.
 24
 25This dataset is from the publication Kraemer et al. (2026), see the Zenodo records for the details.
 26Please cite it if you use this dataset for your research.
 27"""
 28
 29import io
 30import os
 31import json
 32import uuid
 33import zlib
 34import struct
 35from glob import glob
 36from natsort import natsorted
 37from typing import Union, Tuple, Optional, List, Sequence
 38from concurrent import futures
 39
 40import numpy as np
 41from tqdm import tqdm
 42
 43from torch.utils.data import Dataset, DataLoader
 44
 45import torch_em
 46
 47from .. import util
 48
 49
 50URL_BASE = "https://zenodo.org/api/records"
 51SEGMENTATION_RECORD = 20391579
 52
 53DATASETS = {
 54    "DS0002": (16101265, "Kraemer2026A-DS0002-Part2-F-D.zip", [5]),
 55    "DS0004": (16104477, "Kraemer2026A-DS0004-Part2-F-D.zip", [1]),
 56    "DS0005": (16110531, "Kraemer2026A-DS0005-Part2-F-D.zip", [3]),
 57    "DS0006": (16139176, "Kraemer2026A-DS0006-Part2-F-D.zip", [2, 36]),
 58    "DS0007": (16141730, "Kraemer2026A-DS0007-Part2-F-D.zip", [7]),
 59    "DS0008": (16144559, "Kraemer2026A-DS0008-Part2-F-D.zip", [6]),
 60    "DS0009": (16147573, "Kraemer2026A-DS0009-Part2-F-D.zip", [16]),
 61    "DS0010": (16149318, "Kraemer2026A-DS0010-Part2-F-D.zip", [3]),
 62    "DS0011": (16150544, "Kraemer2026A-DS0011-Part2-F-D.zip", [1]),
 63    "DS0012": (16151472, "Kraemer2026A-DS0012-Part2-F-D.zip", [2]),
 64    "DS0013": (16153432, "Kraemer2026A-DS0013-Part2-F-D.zip", [9]),
 65    "DS0014": (16158358, "Kraemer2026A-DS0014-Part2-F-D.zip", [2]),
 66    "DS0015": (16159538, "Kraemer2026A-DS0015-Part2-F-D.zip", [9]),
 67    "DS0016": (16161996, "Kraemer2026A-DS0016-Part2-F-D.zip", [1]),
 68}
 69"""Mapping from the name of a dataset to the Zenodo record and file with its raw data and to its labeled time points."""
 70
 71UNASSIGNED_ID = 65535
 72
 73
 74def _read_zip_entries(url, cache_path):
 75    import requests
 76
 77    if os.path.exists(cache_path):
 78        with open(cache_path) as f:
 79            return json.load(f)
 80
 81    size = int(requests.head(url, allow_redirects=True).headers["Content-Length"])
 82    tail = requests.get(url, headers={"Range": f"bytes={size - 65536}-{size - 1}"}).content
 83    eocd = tail.rfind(b"PK\x05\x06")
 84    n_entries, cd_size, cd_offset = struct.unpack("<HII", tail[eocd + 10:eocd + 20])
 85    if n_entries == 0xFFFF or cd_offset == 0xFFFFFFFF:
 86        locator = tail.rfind(b"PK\x06\x07")
 87        zip64_offset = struct.unpack("<Q", tail[locator + 8:locator + 16])[0]
 88        zip64 = requests.get(url, headers={"Range": f"bytes={zip64_offset}-{zip64_offset + 55}"}).content
 89        n_entries, cd_size, cd_offset = struct.unpack("<QQQ", zip64[32:56])
 90    directory = requests.get(url, headers={"Range": f"bytes={cd_offset}-{cd_offset + cd_size - 1}"}).content
 91
 92    entries, pos = {}, 0
 93    for _ in range(n_entries):
 94        fields = struct.unpack("<IHHHHHHIIIHHHHHII", directory[pos:pos + 46])
 95        method, crc, csize, usize = fields[4], fields[7], fields[8], fields[9]
 96        name_len, extra_len, comment_len = fields[10], fields[11], fields[12]
 97        header_offset = fields[16]
 98        name = directory[pos + 46:pos + 46 + name_len].decode()
 99        extra = directory[pos + 46 + name_len:pos + 46 + name_len + extra_len]
100        offset = 0
101        while offset < len(extra):
102            tag, field_size = struct.unpack("<HH", extra[offset:offset + 4])
103            if tag == 1:
104                field, field_pos = extra[offset + 4:offset + 4 + field_size], 0
105                if usize == 0xFFFFFFFF:
106                    usize, field_pos = struct.unpack("<Q", field[field_pos:field_pos + 8])[0], field_pos + 8
107                if csize == 0xFFFFFFFF:
108                    csize, field_pos = struct.unpack("<Q", field[field_pos:field_pos + 8])[0], field_pos + 8
109                if header_offset == 0xFFFFFFFF:
110                    header_offset = struct.unpack("<Q", field[field_pos:field_pos + 8])[0]
111            offset += 4 + field_size
112        entries[name] = {"method": method, "crc": crc, "compressed_size": csize, "header_offset": header_offset}
113        pos += 46 + name_len + extra_len + comment_len
114
115    os.makedirs(os.path.dirname(cache_path), exist_ok=True)
116    tmp_path = f"{cache_path}.{uuid.uuid4().hex}.incomplete"
117    with open(tmp_path, "w") as f:
118        json.dump(entries, f)
119    os.replace(tmp_path, cache_path)
120    return entries
121
122
123def _read_zip_member(url, entry):
124    import requests
125
126    offset = entry["header_offset"]
127    header = requests.get(url, headers={"Range": f"bytes={offset}-{offset + 29}"}).content
128    name_len, extra_len = struct.unpack("<HH", header[26:30])
129    start = offset + 30 + name_len + extra_len
130    response = requests.get(url, headers={"Range": f"bytes={start}-{start + entry['compressed_size'] - 1}"})
131    response.raise_for_status()
132    data = zlib.decompress(response.content, -15) if entry["method"] == 8 else response.content
133    if zlib.crc32(data) != entry["crc"]:
134        raise RuntimeError("The CRC32 of a downloaded archive member does not match, please try again.")
135    return data
136
137
138def _convert_volume(name, time_point, path, raw_entries, segmentation_entries):
139    import h5py
140    import tifffile
141
142    out_path = os.path.join(path, "preprocessed", f"{name}_TP{time_point:04d}.h5")
143    if os.path.exists(out_path):
144        return
145
146    record, raw_file, _ = DATASETS[name]
147    stem = f"Kraemer2026A-{name}TP{time_point:04d}DR(D1)CH0001PL"
148    raw_member = f"(D1)-ZStacks-Decon-ZS/CH0001/DR(D1)/{stem}(ZS).TIF"
149    label_member = f"Kraemer2026A-{name}-Segmentation/{stem}(YD).TIF"
150    raw_url = f"{URL_BASE}/{record}/files/{raw_file}/content"
151    label_url = f"{URL_BASE}/{SEGMENTATION_RECORD}/files/Kraemer2026A-{name}-Segmentation.zip/content"
152
153    raw = tifffile.imread(io.BytesIO(_read_zip_member(raw_url, raw_entries[raw_member])))
154    labels = tifffile.imread(io.BytesIO(_read_zip_member(label_url, segmentation_entries[label_member])))
155    labels = np.flip(labels.transpose(1, 0, 2), 0)
156    assert raw.shape == labels.shape, f"{name} TP{time_point}: {raw.shape} != {labels.shape}"
157
158    labels = np.where(labels == UNASSIGNED_ID, 0, labels)
159    present = np.unique(labels)
160    present = present[present > 0]
161    lut = np.zeros(UNASSIGNED_ID + 1, dtype="uint16")
162    lut[present] = np.arange(1, len(present) + 1)
163    labels = lut[labels]
164
165    chunks = (64, 128, 128)
166    tmp_path = f"{out_path}.{uuid.uuid4().hex}.incomplete"
167    with h5py.File(tmp_path, "w") as f:
168        f.create_dataset("raw", data=raw, chunks=chunks, compression="gzip")
169        f.create_dataset("labels", data=labels, chunks=chunks, compression="gzip")
170    os.replace(tmp_path, out_path)
171
172
173def _validate(datasets):
174    datasets = list(DATASETS) if datasets is None else list(datasets)
175    invalid = [name for name in datasets if name not in DATASETS]
176    if invalid:
177        raise ValueError(f"{invalid} are not valid datasets. Choose from {list(DATASETS)}.")
178    return datasets
179
180
181def get_slice2_data(
182    path: Union[os.PathLike, str],
183    datasets: Optional[Sequence[str]] = None,
184    n_workers: int = 2,
185    download: bool = False,
186) -> str:
187    """Download the SLICE-2 dataset and convert the labeled volumes to hdf5 files.
188
189    NOTE: Each volume takes about 0.55 GB (raw) to download. Use `datasets` to only download a subset.
190
191    Args:
192        path: Filepath to a folder where the data is downloaded for further processing.
193        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
194        n_workers: The number of parallel download and conversion workers. A worker needs about 3 GB of memory.
195        download: Whether to download the data if it is not present.
196
197    Returns:
198        Filepath where the converted data is stored.
199    """
200    datasets = _validate(datasets)
201    preprocessed_dir = os.path.join(path, "preprocessed")
202    pending = [
203        (name, time_point) for name in datasets for time_point in DATASETS[name][2]
204        if not os.path.exists(os.path.join(preprocessed_dir, f"{name}_TP{time_point:04d}.h5"))
205    ]
206    if not pending:
207        return preprocessed_dir
208    if not download:
209        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
210
211    os.makedirs(preprocessed_dir, exist_ok=True)
212    entries_dir = os.path.join(path, "entries")
213    raw_entries, segmentation_entries = {}, {}
214    for name in sorted({name for name, _ in pending}):
215        record, raw_file, _ = DATASETS[name]
216        raw_entries[name] = _read_zip_entries(
217            f"{URL_BASE}/{record}/files/{raw_file}/content", os.path.join(entries_dir, f"{name}_raw.json")
218        )
219        segmentation_entries[name] = _read_zip_entries(
220            f"{URL_BASE}/{SEGMENTATION_RECORD}/files/Kraemer2026A-{name}-Segmentation.zip/content",
221            os.path.join(entries_dir, f"{name}_segmentation.json"),
222        )
223
224    with futures.ThreadPoolExecutor(n_workers) as pool:
225        tasks = [
226            pool.submit(_convert_volume, name, time_point, path, raw_entries[name], segmentation_entries[name])
227            for name, time_point in pending
228        ]
229        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Download SLICE-2 volumes"):
230            task.result()
231
232    return preprocessed_dir
233
234
235def get_slice2_paths(
236    path: Union[os.PathLike, str],
237    datasets: Optional[Sequence[str]] = None,
238    download: bool = False,
239) -> List[str]:
240    """Get paths to the SLICE-2 data.
241
242    Args:
243        path: Filepath to a folder where the data is downloaded for further processing.
244        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
245        download: Whether to download the data if it is not present.
246
247    Returns:
248        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
249    """
250    datasets = _validate(datasets)
251    data_dir = get_slice2_data(path, datasets, download=download)
252    data_paths = natsorted([
253        p for p in glob(os.path.join(data_dir, "*.h5")) if os.path.basename(p).split("_")[0] in datasets
254    ])
255    assert len(data_paths) > 0
256    return data_paths
257
258
259def get_slice2_dataset(
260    path: Union[os.PathLike, str],
261    patch_shape: Tuple[int, int, int],
262    datasets: Optional[Sequence[str]] = None,
263    offsets: Optional[List[List[int]]] = None,
264    boundaries: bool = False,
265    binary: bool = False,
266    download: bool = False,
267    **kwargs
268) -> Dataset:
269    """Get the SLICE-2 dataset for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.
270
271    Args:
272        path: Filepath to a folder where the data is downloaded for further processing.
273        patch_shape: The patch shape to use for training.
274        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
275        offsets: Offset values for affinity computation used as target.
276        boundaries: Whether to compute boundaries as the target.
277        binary: Whether to use a binary segmentation target.
278        download: Whether to download the data if it is not present.
279        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
280
281    Returns:
282        The segmentation dataset.
283    """
284    data_paths = get_slice2_paths(path, datasets, download)
285
286    kwargs = util.ensure_transforms(ndim=3, **kwargs)
287    kwargs, _ = util.add_instance_label_transform(
288        kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary
289    )
290
291    return torch_em.default_segmentation_dataset(
292        raw_paths=data_paths,
293        raw_key="raw",
294        label_paths=data_paths,
295        label_key="labels",
296        patch_shape=patch_shape,
297        ndim=3,
298        **kwargs
299    )
300
301
302def get_slice2_loader(
303    path: Union[os.PathLike, str],
304    batch_size: int,
305    patch_shape: Tuple[int, int, int],
306    datasets: Optional[Sequence[str]] = None,
307    offsets: Optional[List[List[int]]] = None,
308    boundaries: bool = False,
309    binary: bool = False,
310    download: bool = False,
311    **kwargs
312) -> DataLoader:
313    """Get the SLICE-2 dataloader for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.
314
315    Args:
316        path: Filepath to a folder where the data is downloaded for further processing.
317        batch_size: The batch size for training.
318        patch_shape: The patch shape to use for training.
319        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
320        offsets: Offset values for affinity computation used as target.
321        boundaries: Whether to compute boundaries as the target.
322        binary: Whether to use a binary segmentation target.
323        download: Whether to download the data if it is not present.
324        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
325
326    Returns:
327        The DataLoader.
328    """
329    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
330    dataset = get_slice2_dataset(
331        path, patch_shape, datasets, offsets=offsets, boundaries=boundaries, binary=binary, download=download,
332        **ds_kwargs
333    )
334    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL_BASE = 'https://zenodo.org/api/records'
SEGMENTATION_RECORD = 20391579
DATASETS = {'DS0002': (16101265, 'Kraemer2026A-DS0002-Part2-F-D.zip', [5]), 'DS0004': (16104477, 'Kraemer2026A-DS0004-Part2-F-D.zip', [1]), 'DS0005': (16110531, 'Kraemer2026A-DS0005-Part2-F-D.zip', [3]), 'DS0006': (16139176, 'Kraemer2026A-DS0006-Part2-F-D.zip', [2, 36]), 'DS0007': (16141730, 'Kraemer2026A-DS0007-Part2-F-D.zip', [7]), 'DS0008': (16144559, 'Kraemer2026A-DS0008-Part2-F-D.zip', [6]), 'DS0009': (16147573, 'Kraemer2026A-DS0009-Part2-F-D.zip', [16]), 'DS0010': (16149318, 'Kraemer2026A-DS0010-Part2-F-D.zip', [3]), 'DS0011': (16150544, 'Kraemer2026A-DS0011-Part2-F-D.zip', [1]), 'DS0012': (16151472, 'Kraemer2026A-DS0012-Part2-F-D.zip', [2]), 'DS0013': (16153432, 'Kraemer2026A-DS0013-Part2-F-D.zip', [9]), 'DS0014': (16158358, 'Kraemer2026A-DS0014-Part2-F-D.zip', [2]), 'DS0015': (16159538, 'Kraemer2026A-DS0015-Part2-F-D.zip', [9]), 'DS0016': (16161996, 'Kraemer2026A-DS0016-Part2-F-D.zip', [1])}

Mapping from the name of a dataset to the Zenodo record and file with its raw data and to its labeled time points.

UNASSIGNED_ID = 65535
def get_slice2_data( path: Union[os.PathLike, str], datasets: Optional[Sequence[str]] = None, n_workers: int = 2, download: bool = False) -> str:
182def get_slice2_data(
183    path: Union[os.PathLike, str],
184    datasets: Optional[Sequence[str]] = None,
185    n_workers: int = 2,
186    download: bool = False,
187) -> str:
188    """Download the SLICE-2 dataset and convert the labeled volumes to hdf5 files.
189
190    NOTE: Each volume takes about 0.55 GB (raw) to download. Use `datasets` to only download a subset.
191
192    Args:
193        path: Filepath to a folder where the data is downloaded for further processing.
194        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
195        n_workers: The number of parallel download and conversion workers. A worker needs about 3 GB of memory.
196        download: Whether to download the data if it is not present.
197
198    Returns:
199        Filepath where the converted data is stored.
200    """
201    datasets = _validate(datasets)
202    preprocessed_dir = os.path.join(path, "preprocessed")
203    pending = [
204        (name, time_point) for name in datasets for time_point in DATASETS[name][2]
205        if not os.path.exists(os.path.join(preprocessed_dir, f"{name}_TP{time_point:04d}.h5"))
206    ]
207    if not pending:
208        return preprocessed_dir
209    if not download:
210        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
211
212    os.makedirs(preprocessed_dir, exist_ok=True)
213    entries_dir = os.path.join(path, "entries")
214    raw_entries, segmentation_entries = {}, {}
215    for name in sorted({name for name, _ in pending}):
216        record, raw_file, _ = DATASETS[name]
217        raw_entries[name] = _read_zip_entries(
218            f"{URL_BASE}/{record}/files/{raw_file}/content", os.path.join(entries_dir, f"{name}_raw.json")
219        )
220        segmentation_entries[name] = _read_zip_entries(
221            f"{URL_BASE}/{SEGMENTATION_RECORD}/files/Kraemer2026A-{name}-Segmentation.zip/content",
222            os.path.join(entries_dir, f"{name}_segmentation.json"),
223        )
224
225    with futures.ThreadPoolExecutor(n_workers) as pool:
226        tasks = [
227            pool.submit(_convert_volume, name, time_point, path, raw_entries[name], segmentation_entries[name])
228            for name, time_point in pending
229        ]
230        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Download SLICE-2 volumes"):
231            task.result()
232
233    return preprocessed_dir

Download the SLICE-2 dataset and convert the labeled volumes to hdf5 files.

NOTE: Each volume takes about 0.55 GB (raw) to download. Use datasets to only download a subset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • datasets: The names of the datasets to use, see DATASETS. By default all 14 are used.
  • n_workers: The number of parallel download and conversion workers. A worker needs about 3 GB of memory.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the converted data is stored.

def get_slice2_paths( path: Union[os.PathLike, str], datasets: Optional[Sequence[str]] = None, download: bool = False) -> List[str]:
236def get_slice2_paths(
237    path: Union[os.PathLike, str],
238    datasets: Optional[Sequence[str]] = None,
239    download: bool = False,
240) -> List[str]:
241    """Get paths to the SLICE-2 data.
242
243    Args:
244        path: Filepath to a folder where the data is downloaded for further processing.
245        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
246        download: Whether to download the data if it is not present.
247
248    Returns:
249        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
250    """
251    datasets = _validate(datasets)
252    data_dir = get_slice2_data(path, datasets, download=download)
253    data_paths = natsorted([
254        p for p in glob(os.path.join(data_dir, "*.h5")) if os.path.basename(p).split("_")[0] in datasets
255    ])
256    assert len(data_paths) > 0
257    return data_paths

Get paths to the SLICE-2 data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • datasets: The names of the datasets to use, see DATASETS. By default all 14 are used.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').

def get_slice2_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], datasets: Optional[Sequence[str]] = None, offsets: Optional[List[List[int]]] = None, boundaries: bool = False, binary: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
260def get_slice2_dataset(
261    path: Union[os.PathLike, str],
262    patch_shape: Tuple[int, int, int],
263    datasets: Optional[Sequence[str]] = None,
264    offsets: Optional[List[List[int]]] = None,
265    boundaries: bool = False,
266    binary: bool = False,
267    download: bool = False,
268    **kwargs
269) -> Dataset:
270    """Get the SLICE-2 dataset for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.
271
272    Args:
273        path: Filepath to a folder where the data is downloaded for further processing.
274        patch_shape: The patch shape to use for training.
275        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
276        offsets: Offset values for affinity computation used as target.
277        boundaries: Whether to compute boundaries as the target.
278        binary: Whether to use a binary segmentation target.
279        download: Whether to download the data if it is not present.
280        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
281
282    Returns:
283        The segmentation dataset.
284    """
285    data_paths = get_slice2_paths(path, datasets, download)
286
287    kwargs = util.ensure_transforms(ndim=3, **kwargs)
288    kwargs, _ = util.add_instance_label_transform(
289        kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary
290    )
291
292    return torch_em.default_segmentation_dataset(
293        raw_paths=data_paths,
294        raw_key="raw",
295        label_paths=data_paths,
296        label_key="labels",
297        patch_shape=patch_shape,
298        ndim=3,
299        **kwargs
300    )

Get the SLICE-2 dataset for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • datasets: The names of the datasets to use, see DATASETS. By default all 14 are used.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • binary: Whether to use a binary segmentation target.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_slice2_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int, int], datasets: Optional[Sequence[str]] = None, offsets: Optional[List[List[int]]] = None, boundaries: bool = False, binary: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
303def get_slice2_loader(
304    path: Union[os.PathLike, str],
305    batch_size: int,
306    patch_shape: Tuple[int, int, int],
307    datasets: Optional[Sequence[str]] = None,
308    offsets: Optional[List[List[int]]] = None,
309    boundaries: bool = False,
310    binary: bool = False,
311    download: bool = False,
312    **kwargs
313) -> DataLoader:
314    """Get the SLICE-2 dataloader for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.
315
316    Args:
317        path: Filepath to a folder where the data is downloaded for further processing.
318        batch_size: The batch size for training.
319        patch_shape: The patch shape to use for training.
320        datasets: The names of the datasets to use, see `DATASETS`. By default all 14 are used.
321        offsets: Offset values for affinity computation used as target.
322        boundaries: Whether to compute boundaries as the target.
323        binary: Whether to use a binary segmentation target.
324        download: Whether to download the data if it is not present.
325        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
326
327    Returns:
328        The DataLoader.
329    """
330    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
331    dataset = get_slice2_dataset(
332        path, patch_shape, datasets, offsets=offsets, boundaries=boundaries, binary=binary, download=download,
333        **ds_kwargs
334    )
335    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the SLICE-2 dataloader for 3D nucleus segmentation in light-sheet microscopy of Tribolium embryos.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • datasets: The names of the datasets to use, see DATASETS. By default all 14 are used.
  • offsets: Offset values for affinity computation used as target.
  • boundaries: Whether to compute boundaries as the target.
  • binary: Whether to use a binary segmentation target.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.