torch_em.data.datasets.histopathology.ecm_phenotyping
This dataset contains cell instance segmentation annotations for imaging mass cytometry (IMC) of formalin-fixed paraffin-embedded (FFPE) mouse lung sections, imaged to study the relationship between cellular organisation and extracellular matrix (ECM) composition during allergic airway inflammation.
The data is from the publication "Extracellular matrix phenotyping by imaging mass cytometry defines distinct cellular matrix environments associated with allergic airway inflammation" and hosted on the BioImage Archive at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3184. It is available under the CC BY 4.0 license. Please cite it if you use this dataset in your research.
This loader covers the "Hyperion Imaging of Stained Allergic Mouse Lung" study component, which provides 36 acquisitions (3 slides, 12 ROIs each) as processed multichannel TIFF stacks (76 channels, covering a 42-marker antibody panel plus derived distance-transform channels) with matching per-cell instance segmentation masks generated with DeepCell. Note that, despite the "ECM phenotyping" framing of the study, the ECM channels themselves are only provided as continuous per-pixel distance-transform maps (see the "dist" folder on the BioImage Archive), not as discrete class-labeled ECM masks; the only raster segmentation masks in this study are the generic per-cell instance masks used here as the label target. The study also hosts two further "removable" components (immunofluorescence slide scans and confocal imaging of precision cut lung slices) with only geometrical (QuPath / arivis) region annotations rather than instance or semantic masks; these are out of scope for this loader.
1"""This dataset contains cell instance segmentation annotations for imaging mass cytometry (IMC) 2of formalin-fixed paraffin-embedded (FFPE) mouse lung sections, imaged to study the relationship 3between cellular organisation and extracellular matrix (ECM) composition during allergic airway 4inflammation. 5 6The data is from the publication "Extracellular matrix phenotyping by imaging mass cytometry 7defines distinct cellular matrix environments associated with allergic airway inflammation" and 8hosted on the BioImage Archive at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3184. 9It is available under the CC BY 4.0 license. Please cite it if you use this dataset in your 10research. 11 12This loader covers the "Hyperion Imaging of Stained Allergic Mouse Lung" study component, which 13provides 36 acquisitions (3 slides, 12 ROIs each) as processed multichannel TIFF stacks (76 14channels, covering a 42-marker antibody panel plus derived distance-transform channels) with 15matching per-cell instance segmentation masks generated with DeepCell. Note that, despite the 16"ECM phenotyping" framing of the study, the ECM channels themselves are only provided as 17continuous per-pixel distance-transform maps (see the "dist" folder on the BioImage Archive), 18not as discrete class-labeled ECM masks; the only raster segmentation masks in this study are 19the generic per-cell instance masks used here as the label target. The study also hosts two 20further "removable" components (immunofluorescence slide scans and confocal imaging of precision 21cut lung slices) with only geometrical (QuPath / arivis) region annotations rather than instance 22or semantic masks; these are out of scope for this loader. 23""" 24 25import os 26from glob import glob 27from typing import List, Optional, Sequence, Tuple, Union 28 29from torch.utils.data import Dataset, DataLoader 30 31import torch_em 32 33from .. import util 34 35 36BASE_URL = "https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/184/S-BIAD3184/Files/data_upload/processed_data" 37 38SLIDES = ("slide1", "slide2", "slide3") 39ROIS = tuple(f"{i:03d}" for i in range(1, 13)) 40 41 42def get_ecm_phenotyping_data( 43 path: Union[os.PathLike, str], 44 slides: Optional[Sequence[str]] = None, 45 download: bool = False, 46) -> str: 47 """Download the ECM phenotyping (allergic mouse lung IMC) data. 48 49 Args: 50 path: Filepath to a folder where the downloaded data will be saved. 51 slides: The slide names to prepare. By default all three slides (36 acquisitions in 52 total) are prepared, which requires downloading several gigabytes of data. 53 download: Whether to download the data if it is not present. 54 55 Returns: 56 Filepath to the folder where the data is stored. 57 """ 58 if slides is None: 59 slides = SLIDES 60 else: 61 invalid = sorted(set(slides) - set(SLIDES)) 62 if invalid: 63 raise ValueError(f"Invalid slide name(s) {invalid}. Choose from {SLIDES}.") 64 65 raw_dir = os.path.join(path, "img") 66 label_dir = os.path.join(path, "masks") 67 os.makedirs(raw_dir, exist_ok=True) 68 os.makedirs(label_dir, exist_ok=True) 69 70 for slide in slides: 71 for roi in ROIS: 72 fname = f"m16263_{slide}_{roi}.tiff" 73 raw_path = os.path.join(raw_dir, fname) 74 label_path = os.path.join(label_dir, fname) 75 util.download_source(raw_path, f"{BASE_URL}/img/{fname}", download, checksum=None) 76 util.download_source(label_path, f"{BASE_URL}/masks/{fname}", download, checksum=None) 77 78 return path 79 80 81def get_ecm_phenotyping_paths( 82 path: Union[os.PathLike, str], 83 slides: Optional[Sequence[str]] = None, 84 download: bool = False, 85) -> Tuple[List[str], List[str]]: 86 """Get paths to the ECM phenotyping images and cell instance segmentation masks. 87 88 Args: 89 path: Filepath to a folder where the downloaded data will be saved. 90 slides: The slide names to load. By default all three slides are loaded. 91 download: Whether to download the data if it is not present. 92 93 Returns: 94 The image paths and corresponding mask paths. 95 """ 96 root = get_ecm_phenotyping_data(path, slides, download) 97 98 raw_paths = sorted(glob(os.path.join(root, "img", "*.tiff"))) 99 label_paths = sorted(glob(os.path.join(root, "masks", "*.tiff"))) 100 101 missing_paths = [p for p in raw_paths + label_paths if not os.path.exists(p)] 102 if missing_paths: 103 raise RuntimeError(f"Could not find {len(missing_paths)} ECM phenotyping files.") 104 105 return raw_paths, label_paths 106 107 108def get_ecm_phenotyping_dataset( 109 path: Union[os.PathLike, str], 110 patch_shape: Tuple[int, int], 111 slides: Optional[Sequence[str]] = None, 112 download: bool = False, 113 resize_inputs: bool = False, 114 **kwargs 115) -> Dataset: 116 """Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry. 117 118 Args: 119 path: Filepath to a folder where the downloaded data will be saved. 120 patch_shape: The 2D patch shape to use for training. 121 slides: The slide names to load. By default all three slides are loaded. 122 download: Whether to download the data if it is not present. 123 resize_inputs: Whether to resize the input images. 124 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 125 126 Returns: 127 The segmentation dataset. 128 """ 129 if len(patch_shape) != 2: 130 raise ValueError(f"The ECM phenotyping patch shape must be two-dimensional, got {patch_shape}.") 131 132 raw_paths, label_paths = get_ecm_phenotyping_paths(path, slides, download) 133 134 if resize_inputs: 135 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 136 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 137 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 138 ) 139 140 return torch_em.default_segmentation_dataset( 141 raw_paths=raw_paths, 142 raw_key=None, 143 label_paths=label_paths, 144 label_key=None, 145 patch_shape=patch_shape, 146 is_seg_dataset=True, 147 with_channels=True, 148 ndim=2, 149 **kwargs 150 ) 151 152 153def get_ecm_phenotyping_loader( 154 path: Union[os.PathLike, str], 155 patch_shape: Tuple[int, int], 156 batch_size: int, 157 slides: Optional[Sequence[str]] = None, 158 download: bool = False, 159 resize_inputs: bool = False, 160 **kwargs 161) -> DataLoader: 162 """Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry. 163 164 Args: 165 path: Filepath to a folder where the downloaded data will be saved. 166 patch_shape: The 2D patch shape to use for training. 167 batch_size: The batch size for training. 168 slides: The slide names to load. By default all three slides are loaded. 169 download: Whether to download the data if it is not present. 170 resize_inputs: Whether to resize the input images. 171 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for 172 the PyTorch DataLoader. 173 174 Returns: 175 The DataLoader. 176 """ 177 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 178 dataset = get_ecm_phenotyping_dataset( 179 path, patch_shape, slides=slides, download=download, resize_inputs=resize_inputs, **ds_kwargs 180 ) 181 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
43def get_ecm_phenotyping_data( 44 path: Union[os.PathLike, str], 45 slides: Optional[Sequence[str]] = None, 46 download: bool = False, 47) -> str: 48 """Download the ECM phenotyping (allergic mouse lung IMC) data. 49 50 Args: 51 path: Filepath to a folder where the downloaded data will be saved. 52 slides: The slide names to prepare. By default all three slides (36 acquisitions in 53 total) are prepared, which requires downloading several gigabytes of data. 54 download: Whether to download the data if it is not present. 55 56 Returns: 57 Filepath to the folder where the data is stored. 58 """ 59 if slides is None: 60 slides = SLIDES 61 else: 62 invalid = sorted(set(slides) - set(SLIDES)) 63 if invalid: 64 raise ValueError(f"Invalid slide name(s) {invalid}. Choose from {SLIDES}.") 65 66 raw_dir = os.path.join(path, "img") 67 label_dir = os.path.join(path, "masks") 68 os.makedirs(raw_dir, exist_ok=True) 69 os.makedirs(label_dir, exist_ok=True) 70 71 for slide in slides: 72 for roi in ROIS: 73 fname = f"m16263_{slide}_{roi}.tiff" 74 raw_path = os.path.join(raw_dir, fname) 75 label_path = os.path.join(label_dir, fname) 76 util.download_source(raw_path, f"{BASE_URL}/img/{fname}", download, checksum=None) 77 util.download_source(label_path, f"{BASE_URL}/masks/{fname}", download, checksum=None) 78 79 return path
Download the ECM phenotyping (allergic mouse lung IMC) data.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- slides: The slide names to prepare. By default all three slides (36 acquisitions in total) are prepared, which requires downloading several gigabytes of data.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the data is stored.
82def get_ecm_phenotyping_paths( 83 path: Union[os.PathLike, str], 84 slides: Optional[Sequence[str]] = None, 85 download: bool = False, 86) -> Tuple[List[str], List[str]]: 87 """Get paths to the ECM phenotyping images and cell instance segmentation masks. 88 89 Args: 90 path: Filepath to a folder where the downloaded data will be saved. 91 slides: The slide names to load. By default all three slides are loaded. 92 download: Whether to download the data if it is not present. 93 94 Returns: 95 The image paths and corresponding mask paths. 96 """ 97 root = get_ecm_phenotyping_data(path, slides, download) 98 99 raw_paths = sorted(glob(os.path.join(root, "img", "*.tiff"))) 100 label_paths = sorted(glob(os.path.join(root, "masks", "*.tiff"))) 101 102 missing_paths = [p for p in raw_paths + label_paths if not os.path.exists(p)] 103 if missing_paths: 104 raise RuntimeError(f"Could not find {len(missing_paths)} ECM phenotyping files.") 105 106 return raw_paths, label_paths
Get paths to the ECM phenotyping images and cell instance segmentation masks.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- slides: The slide names to load. By default all three slides are loaded.
- download: Whether to download the data if it is not present.
Returns:
The image paths and corresponding mask paths.
109def get_ecm_phenotyping_dataset( 110 path: Union[os.PathLike, str], 111 patch_shape: Tuple[int, int], 112 slides: Optional[Sequence[str]] = None, 113 download: bool = False, 114 resize_inputs: bool = False, 115 **kwargs 116) -> Dataset: 117 """Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry. 118 119 Args: 120 path: Filepath to a folder where the downloaded data will be saved. 121 patch_shape: The 2D patch shape to use for training. 122 slides: The slide names to load. By default all three slides are loaded. 123 download: Whether to download the data if it is not present. 124 resize_inputs: Whether to resize the input images. 125 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 126 127 Returns: 128 The segmentation dataset. 129 """ 130 if len(patch_shape) != 2: 131 raise ValueError(f"The ECM phenotyping patch shape must be two-dimensional, got {patch_shape}.") 132 133 raw_paths, label_paths = get_ecm_phenotyping_paths(path, slides, download) 134 135 if resize_inputs: 136 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 137 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 138 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 139 ) 140 141 return torch_em.default_segmentation_dataset( 142 raw_paths=raw_paths, 143 raw_key=None, 144 label_paths=label_paths, 145 label_key=None, 146 patch_shape=patch_shape, 147 is_seg_dataset=True, 148 with_channels=True, 149 ndim=2, 150 **kwargs 151 )
Get the ECM phenotyping dataset for cell instance segmentation in imaging mass cytometry.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- patch_shape: The 2D patch shape to use for training.
- slides: The slide names to load. By default all three slides are loaded.
- download: Whether to download the data if it is not present.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
154def get_ecm_phenotyping_loader( 155 path: Union[os.PathLike, str], 156 patch_shape: Tuple[int, int], 157 batch_size: int, 158 slides: Optional[Sequence[str]] = None, 159 download: bool = False, 160 resize_inputs: bool = False, 161 **kwargs 162) -> DataLoader: 163 """Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry. 164 165 Args: 166 path: Filepath to a folder where the downloaded data will be saved. 167 patch_shape: The 2D patch shape to use for training. 168 batch_size: The batch size for training. 169 slides: The slide names to load. By default all three slides are loaded. 170 download: Whether to download the data if it is not present. 171 resize_inputs: Whether to resize the input images. 172 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for 173 the PyTorch DataLoader. 174 175 Returns: 176 The DataLoader. 177 """ 178 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 179 dataset = get_ecm_phenotyping_dataset( 180 path, patch_shape, slides=slides, download=download, resize_inputs=resize_inputs, **ds_kwargs 181 ) 182 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the ECM phenotyping dataloader for cell instance segmentation in imaging mass cytometry.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- patch_shape: The 2D patch shape to use for training.
- batch_size: The batch size for training.
- slides: The slide names to load. By default all three slides are loaded.
- download: Whether to download the data if it is not present.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.