torch_em.data.datasets.histopathology.ecm_phenotyping

This dataset contains cell instance segmentation annotations for imaging mass cytometry (IMC) of formalin-fixed paraffin-embedded (FFPE) mouse lung sections, imaged to study the relationship between cellular organisation and extracellular matrix (ECM) composition during allergic airway inflammation.

The data is from the publication "Extracellular matrix phenotyping by imaging mass cytometry defines distinct cellular matrix environments associated with allergic airway inflammation" and hosted on the BioImage Archive at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3184. It is available under the CC BY 4.0 license. Please cite it if you use this dataset in your research.

This loader covers the "Hyperion Imaging of Stained Allergic Mouse Lung" study component, which provides 36 acquisitions (3 slides, 12 ROIs each) as processed multichannel TIFF stacks (76 channels, covering a 42-marker antibody panel plus derived distance-transform channels) with matching per-cell instance segmentation masks generated with DeepCell. Note that, despite the "ECM phenotyping" framing of the study, the ECM channels themselves are only provided as continuous per-pixel distance-transform maps (see the "dist" folder on the BioImage Archive), not as discrete class-labeled ECM masks; the only raster segmentation masks in this study are the generic per-cell instance masks used here as the label target. The study also hosts two further "removable" components (immunofluorescence slide scans and confocal imaging of precision cut lung slices) with only geometrical (QuPath / arivis) region annotations rather than instance or semantic masks; these are out of scope for this loader.

  1"""This dataset contains cell instance segmentation annotations for imaging mass cytometry (IMC)
  2of formalin-fixed paraffin-embedded (FFPE) mouse lung sections, imaged to study the relationship
  3between cellular organisation and extracellular matrix (ECM) composition during allergic airway
  4inflammation.
  5
  6The data is from the publication "Extracellular matrix phenotyping by imaging mass cytometry
  7defines distinct cellular matrix environments associated with allergic airway inflammation" and
  8hosted on the BioImage Archive at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3184.
  9It is available under the CC BY 4.0 license. Please cite it if you use this dataset in your
 10research.
 11
 12This loader covers the "Hyperion Imaging of Stained Allergic Mouse Lung" study component, which
 13provides 36 acquisitions (3 slides, 12 ROIs each) as processed multichannel TIFF stacks (76
 14channels, covering a 42-marker antibody panel plus derived distance-transform channels) with
 15matching per-cell instance segmentation masks generated with DeepCell. Note that, despite the
 16"ECM phenotyping" framing of the study, the ECM channels themselves are only provided as
 17continuous per-pixel distance-transform maps (see the "dist" folder on the BioImage Archive),
 18not as discrete class-labeled ECM masks; the only raster segmentation masks in this study are
 19the generic per-cell instance masks used here as the label target. The study also hosts two
 20further "removable" components (immunofluorescence slide scans and confocal imaging of precision
 21cut lung slices) with only geometrical (QuPath / arivis) region annotations rather than instance
 22or semantic masks; these are out of scope for this loader.
 23"""
 24
 25import os
 26from glob import glob
 27from typing import List, Optional, Sequence, Tuple, Union
 28
 29from torch.utils.data import Dataset, DataLoader
 30
 31import torch_em
 32
 33from .. import util
 34
 35
 36BASE_URL = "https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/184/S-BIAD3184/Files/data_upload/processed_data"
 37
 38SLIDES = ("slide1", "slide2", "slide3")
 39ROIS = tuple(f"{i:03d}" for i in range(1, 13))
 40
 41
 42def get_ecm_phenotyping_data(
 43    path: Union[os.PathLike, str],
 44    slides: Optional[Sequence[str]] = None,
 45    download: bool = False,
 46) -> str:
 47    """Download the ECM phenotyping (allergic mouse lung IMC) data.
 48
 49    Args:
 50        path: Filepath to a folder where the downloaded data will be saved.
 51        slides: The slide names to prepare. By default all three slides (36 acquisitions in
 52            total) are prepared, which requires downloading several gigabytes of data.
 53        download: Whether to download the data if it is not present.
 54
 55    Returns:
 56        Filepath to the folder where the data is stored.
 57    """
 58    if slides is None:
 59        slides = SLIDES
 60    else:
 61        invalid = sorted(set(slides) - set(SLIDES))
 62        if invalid:
 63            raise ValueError(f"Invalid slide name(s) {invalid}. Choose from {SLIDES}.")
 64
 65    raw_dir = os.path.join(path, "img")
 66    label_dir = os.path.join(path, "masks")
 67    os.makedirs(raw_dir, exist_ok=True)
 68    os.makedirs(label_dir, exist_ok=True)
 69
 70    for slide in slides:
 71        for roi in ROIS:
 72            fname = f"m16263_{slide}_{roi}.tiff"
 73            raw_path = os.path.join(raw_dir, fname)
 74            label_path = os.path.join(label_dir, fname)
 75            util.download_source(raw_path, f"{BASE_URL}/img/{fname}", download, checksum=None)
 76            util.download_source(label_path, f"{BASE_URL}/masks/{fname}", download, checksum=None)
 77
 78    return path
 79
 80
 81def get_ecm_phenotyping_paths(
 82    path: Union[os.PathLike, str],
 83    slides: Optional[Sequence[str]] = None,
 84    download: bool = False,
 85) -> Tuple[List[str], List[str]]:
 86    """Get paths to the ECM phenotyping images and cell instance segmentation masks.
 87
 88    Args:
 89        path: Filepath to a folder where the downloaded data will be saved.
 90        slides: The slide names to load. By default all three slides are loaded.
 91        download: Whether to download the data if it is not present.
 92
 93    Returns:
 94        The image paths and corresponding mask paths.
 95    """
 96    root = get_ecm_phenotyping_data(path, slides, download)
 97
 98    raw_paths = sorted(glob(os.path.join(root, "img", "*.tiff")))
 99    label_paths = sorted(glob(os.path.join(root, "masks", "*.tiff")))
100
101    missing_paths = [p for p in raw_paths + label_paths if not os.path.exists(p)]
102    if missing_paths:
103        raise RuntimeError(f"Could not find {len(missing_paths)} ECM phenotyping files.")
104
105    return raw_paths, label_paths
106
107
108def get_ecm_phenotyping_dataset(
109    path: Union[os.PathLike, str],
110    patch_shape: Tuple[int, int],
111    slides: Optional[Sequence[str]] = None,
112    download: bool = False,
113    resize_inputs: bool = False,
114    **kwargs
115) -> Dataset:
116    """Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry.
117
118    Args:
119        path: Filepath to a folder where the downloaded data will be saved.
120        patch_shape: The 2D patch shape to use for training.
121        slides: The slide names to load. By default all three slides are loaded.
122        download: Whether to download the data if it is not present.
123        resize_inputs: Whether to resize the input images.
124        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
125
126    Returns:
127        The segmentation dataset.
128    """
129    if len(patch_shape) != 2:
130        raise ValueError(f"The ECM phenotyping patch shape must be two-dimensional, got {patch_shape}.")
131
132    raw_paths, label_paths = get_ecm_phenotyping_paths(path, slides, download)
133
134    if resize_inputs:
135        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
136        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
137            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
138        )
139
140    return torch_em.default_segmentation_dataset(
141        raw_paths=raw_paths,
142        raw_key=None,
143        label_paths=label_paths,
144        label_key=None,
145        patch_shape=patch_shape,
146        is_seg_dataset=True,
147        with_channels=True,
148        ndim=2,
149        **kwargs
150    )
151
152
153def get_ecm_phenotyping_loader(
154    path: Union[os.PathLike, str],
155    patch_shape: Tuple[int, int],
156    batch_size: int,
157    slides: Optional[Sequence[str]] = None,
158    download: bool = False,
159    resize_inputs: bool = False,
160    **kwargs
161) -> DataLoader:
162    """Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry.
163
164    Args:
165        path: Filepath to a folder where the downloaded data will be saved.
166        patch_shape: The 2D patch shape to use for training.
167        batch_size: The batch size for training.
168        slides: The slide names to load. By default all three slides are loaded.
169        download: Whether to download the data if it is not present.
170        resize_inputs: Whether to resize the input images.
171        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for
172            the PyTorch DataLoader.
173
174    Returns:
175        The DataLoader.
176    """
177    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
178    dataset = get_ecm_phenotyping_dataset(
179        path, patch_shape, slides=slides, download=download, resize_inputs=resize_inputs, **ds_kwargs
180    )
181    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
BASE_URL = 'https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/184/S-BIAD3184/Files/data_upload/processed_data'
SLIDES = ('slide1', 'slide2', 'slide3')
ROIS = ('001', '002', '003', '004', '005', '006', '007', '008', '009', '010', '011', '012')
def get_ecm_phenotyping_data( path: Union[os.PathLike, str], slides: Optional[Sequence[str]] = None, download: bool = False) -> str:
43def get_ecm_phenotyping_data(
44    path: Union[os.PathLike, str],
45    slides: Optional[Sequence[str]] = None,
46    download: bool = False,
47) -> str:
48    """Download the ECM phenotyping (allergic mouse lung IMC) data.
49
50    Args:
51        path: Filepath to a folder where the downloaded data will be saved.
52        slides: The slide names to prepare. By default all three slides (36 acquisitions in
53            total) are prepared, which requires downloading several gigabytes of data.
54        download: Whether to download the data if it is not present.
55
56    Returns:
57        Filepath to the folder where the data is stored.
58    """
59    if slides is None:
60        slides = SLIDES
61    else:
62        invalid = sorted(set(slides) - set(SLIDES))
63        if invalid:
64            raise ValueError(f"Invalid slide name(s) {invalid}. Choose from {SLIDES}.")
65
66    raw_dir = os.path.join(path, "img")
67    label_dir = os.path.join(path, "masks")
68    os.makedirs(raw_dir, exist_ok=True)
69    os.makedirs(label_dir, exist_ok=True)
70
71    for slide in slides:
72        for roi in ROIS:
73            fname = f"m16263_{slide}_{roi}.tiff"
74            raw_path = os.path.join(raw_dir, fname)
75            label_path = os.path.join(label_dir, fname)
76            util.download_source(raw_path, f"{BASE_URL}/img/{fname}", download, checksum=None)
77            util.download_source(label_path, f"{BASE_URL}/masks/{fname}", download, checksum=None)
78
79    return path

Download the ECM phenotyping (allergic mouse lung IMC) data.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • slides: The slide names to prepare. By default all three slides (36 acquisitions in total) are prepared, which requires downloading several gigabytes of data.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the data is stored.

def get_ecm_phenotyping_paths( path: Union[os.PathLike, str], slides: Optional[Sequence[str]] = None, download: bool = False) -> Tuple[List[str], List[str]]:
 82def get_ecm_phenotyping_paths(
 83    path: Union[os.PathLike, str],
 84    slides: Optional[Sequence[str]] = None,
 85    download: bool = False,
 86) -> Tuple[List[str], List[str]]:
 87    """Get paths to the ECM phenotyping images and cell instance segmentation masks.
 88
 89    Args:
 90        path: Filepath to a folder where the downloaded data will be saved.
 91        slides: The slide names to load. By default all three slides are loaded.
 92        download: Whether to download the data if it is not present.
 93
 94    Returns:
 95        The image paths and corresponding mask paths.
 96    """
 97    root = get_ecm_phenotyping_data(path, slides, download)
 98
 99    raw_paths = sorted(glob(os.path.join(root, "img", "*.tiff")))
100    label_paths = sorted(glob(os.path.join(root, "masks", "*.tiff")))
101
102    missing_paths = [p for p in raw_paths + label_paths if not os.path.exists(p)]
103    if missing_paths:
104        raise RuntimeError(f"Could not find {len(missing_paths)} ECM phenotyping files.")
105
106    return raw_paths, label_paths

Get paths to the ECM phenotyping images and cell instance segmentation masks.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • slides: The slide names to load. By default all three slides are loaded.
  • download: Whether to download the data if it is not present.
Returns:

The image paths and corresponding mask paths.

def get_ecm_phenotyping_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], slides: Optional[Sequence[str]] = None, download: bool = False, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
109def get_ecm_phenotyping_dataset(
110    path: Union[os.PathLike, str],
111    patch_shape: Tuple[int, int],
112    slides: Optional[Sequence[str]] = None,
113    download: bool = False,
114    resize_inputs: bool = False,
115    **kwargs
116) -> Dataset:
117    """Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry.
118
119    Args:
120        path: Filepath to a folder where the downloaded data will be saved.
121        patch_shape: The 2D patch shape to use for training.
122        slides: The slide names to load. By default all three slides are loaded.
123        download: Whether to download the data if it is not present.
124        resize_inputs: Whether to resize the input images.
125        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
126
127    Returns:
128        The segmentation dataset.
129    """
130    if len(patch_shape) != 2:
131        raise ValueError(f"The ECM phenotyping patch shape must be two-dimensional, got {patch_shape}.")
132
133    raw_paths, label_paths = get_ecm_phenotyping_paths(path, slides, download)
134
135    if resize_inputs:
136        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
137        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
138            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
139        )
140
141    return torch_em.default_segmentation_dataset(
142        raw_paths=raw_paths,
143        raw_key=None,
144        label_paths=label_paths,
145        label_key=None,
146        patch_shape=patch_shape,
147        is_seg_dataset=True,
148        with_channels=True,
149        ndim=2,
150        **kwargs
151    )

Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • patch_shape: The 2D patch shape to use for training.
  • slides: The slide names to load. By default all three slides are loaded.
  • download: Whether to download the data if it is not present.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_ecm_phenotyping_loader( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], batch_size: int, slides: Optional[Sequence[str]] = None, download: bool = False, resize_inputs: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
154def get_ecm_phenotyping_loader(
155    path: Union[os.PathLike, str],
156    patch_shape: Tuple[int, int],
157    batch_size: int,
158    slides: Optional[Sequence[str]] = None,
159    download: bool = False,
160    resize_inputs: bool = False,
161    **kwargs
162) -> DataLoader:
163    """Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry.
164
165    Args:
166        path: Filepath to a folder where the downloaded data will be saved.
167        patch_shape: The 2D patch shape to use for training.
168        batch_size: The batch size for training.
169        slides: The slide names to load. By default all three slides are loaded.
170        download: Whether to download the data if it is not present.
171        resize_inputs: Whether to resize the input images.
172        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for
173            the PyTorch DataLoader.
174
175    Returns:
176        The DataLoader.
177    """
178    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
179    dataset = get_ecm_phenotyping_dataset(
180        path, patch_shape, slides=slides, download=download, resize_inputs=resize_inputs, **ds_kwargs
181    )
182    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry.

Arguments:
  • path: Filepath to a folder where the downloaded data will be saved.
  • patch_shape: The 2D patch shape to use for training.
  • batch_size: The batch size for training.
  • slides: The slide names to load. By default all three slides are loaded.
  • download: Whether to download the data if it is not present.
  • resize_inputs: Whether to resize the input images.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.