torch_em.data.datasets.medical.inbreast
The INbreast dataset contains annotations for lesion segmentation in full-field digital mammograms.
INbreast was acquired at the Breast Research Group, INESC Porto / Hospital de Sao Joao (Portugal) with a full-field digital mammography (FFDM) system, which distinguishes it from CBIS-DDSM (a re-scan of older film mammograms). It comprises 410 images from 115 cases (90 cases with both breasts, 4 images each; 25 mastectomy cases, 2 images each), each distributed as a DICOM file. Specialists outlined masses, calcifications, calcification clusters, spiculated regions, asymmetries and architectural distortions with the OsiriX viewer, and the resulting contours are exported per case as an OsiriX property list ('.xml', parseable with 'plistlib'). Each ROI lists its pixel coordinates under 'Point_px'; ROIs with 3 or more points are closed polygons (rasterized here with 'skimage.draw.polygon'), ROIs with 1 or 2 points are isolated pixel-level markers (usually individual, non-clustered calcifications).
The label ids used here are (see LABEL_IDS): 0 = background, 1 = mass, 2 = calcification, 3 = cluster
(of calcifications), 4 = spiculated region, 5 = asymmetry, 6 = distortion. The 'Name' field of a few ROIs
in the original xml files is misspelled or inconsistently capitalized (e.g. 'Assymetry', 'Espiculated Region',
'Calcifications'); these are normalized to the canonical names above. A handful of ROIs with an empty or
placeholder name (e.g. 'Unnamed', 'Point 1') carry no lesion information and are skipped.
NOTE: The original INbreast distribution required a request form to the Breast Research Group and is not publicly downloadable anymore. This module downloads the 'INbreast Release 1.0' mirror hosted on Kaggle (https://www.kaggle.com/datasets/ramanathansp20/inbreast-dataset), which reproduces the original release folder structure ('AllDICOMs', 'AllXML', 'AllROI', 'MedicalReports', 'INbreast.xls', 'README.txt') including the per-lesion xml contours. Other INbreast mirrors on Kaggle (e.g. 'martholi/inbreast') only redistribute the DICOM images without the xml annotations and are not suitable for segmentation.
NOTE: Extracting the mirrored zip file with python's 'zipfile' can fail with a 'bad zipfile offset' / 'bad magic number' error, because the archive that Kaggle serves for this dataset has extra bytes prepended to it. This module extracts it with the 'unzip' command line tool instead, which recovers from this offset issue. Please make sure 'unzip' is available (it ships with most Linux distributions).
This dataset is from the publication https://doi.org/10.1016/j.acra.2011.09.014. Please cite it if you use this dataset in your research.
NOTE: The DICOM loading requires the 'pydicom' python package.
1"""The INbreast dataset contains annotations for lesion segmentation in full-field digital mammograms. 2 3INbreast was acquired at the Breast Research Group, INESC Porto / Hospital de Sao Joao (Portugal) with a 4full-field digital mammography (FFDM) system, which distinguishes it from CBIS-DDSM (a re-scan of older 5film mammograms). It comprises 410 images from 115 cases (90 cases with both breasts, 4 images each; 25 6mastectomy cases, 2 images each), each distributed as a DICOM file. Specialists outlined masses, calcifications, 7calcification clusters, spiculated regions, asymmetries and architectural distortions with the OsiriX viewer, 8and the resulting contours are exported per case as an OsiriX property list ('.xml', parseable with 'plistlib'). 9Each ROI lists its pixel coordinates under 'Point_px'; ROIs with 3 or more points are closed polygons 10(rasterized here with 'skimage.draw.polygon'), ROIs with 1 or 2 points are isolated pixel-level markers 11(usually individual, non-clustered calcifications). 12 13The label ids used here are (see `LABEL_IDS`): 0 = background, 1 = mass, 2 = calcification, 3 = cluster 14(of calcifications), 4 = spiculated region, 5 = asymmetry, 6 = distortion. The 'Name' field of a few ROIs 15in the original xml files is misspelled or inconsistently capitalized (e.g. 'Assymetry', 'Espiculated Region', 16'Calcifications'); these are normalized to the canonical names above. A handful of ROIs with an empty or 17placeholder name (e.g. 'Unnamed', 'Point 1') carry no lesion information and are skipped. 18 19NOTE: The original INbreast distribution required a request form to the Breast Research Group and is not 20publicly downloadable anymore. This module downloads the 'INbreast Release 1.0' mirror hosted on Kaggle 21(https://www.kaggle.com/datasets/ramanathansp20/inbreast-dataset), which reproduces the original release 22folder structure ('AllDICOMs', 'AllXML', 'AllROI', 'MedicalReports', 'INbreast.xls', 'README.txt') including 23the per-lesion xml contours. Other INbreast mirrors on Kaggle (e.g. 'martholi/inbreast') only redistribute 24the DICOM images without the xml annotations and are not suitable for segmentation. 25 26NOTE: Extracting the mirrored zip file with python's 'zipfile' can fail with a 'bad zipfile offset' / 27'bad magic number' error, because the archive that Kaggle serves for this dataset has extra bytes prepended 28to it. This module extracts it with the 'unzip' command line tool instead, which recovers from this offset 29issue. Please make sure 'unzip' is available (it ships with most Linux distributions). 30 31This dataset is from the publication https://doi.org/10.1016/j.acra.2011.09.014. Please cite it if you use 32this dataset in your research. 33 34NOTE: The DICOM loading requires the 'pydicom' python package. 35""" 36 37import os 38from glob import glob 39from shutil import which 40from subprocess import run 41from tqdm import tqdm 42from natsort import natsorted 43from typing import Union, Tuple, List 44 45import numpy as np 46 47from torch.utils.data import Dataset, DataLoader 48 49import torch_em 50 51from .. import util 52 53 54LABEL_IDS = { 55 "background": 0, 56 "mass": 1, 57 "calcification": 2, 58 "cluster": 3, 59 "spiculated_region": 4, 60 "asymmetry": 5, 61 "distortion": 6, 62} 63 64# The 'Name' field of the xml ROIs is not fully standardized (typos, capitalization). This maps every 65# variant that was found in the release to a canonical label id. Names that are not listed here (e.g. the 66# empty string, 'Unnamed', 'Point 1') do not carry lesion information and are skipped. 67ROI_NAME_TO_LABEL_ID = { 68 "mass": LABEL_IDS["mass"], 69 "calcification": LABEL_IDS["calcification"], 70 "calcifications": LABEL_IDS["calcification"], 71 "cluster": LABEL_IDS["cluster"], 72 "spiculated region": LABEL_IDS["spiculated_region"], 73 "espiculated region": LABEL_IDS["spiculated_region"], 74 "asymmetry": LABEL_IDS["asymmetry"], 75 "assymetry": LABEL_IDS["asymmetry"], 76 "distortion": LABEL_IDS["distortion"], 77} 78 79 80def _unzip_inbreast(zip_path, dst): 81 # 'zipfile' fails on this archive with a 'bad zipfile offset' error, because Kaggle serves it with 82 # extra bytes prepended. The 'unzip' CLI recovers from this offset issue, so it is used instead of 83 # 'torch_em.data.datasets.util.unzip'. 84 if which("unzip") is None: 85 raise RuntimeError("Need the 'unzip' CLI to extract the INbreast archive.") 86 run(["unzip", "-q", "-o", zip_path, "-d", dst]) 87 os.remove(zip_path) 88 89 90def _parse_xml_rois(xml_path): 91 import plistlib 92 93 with open(xml_path, "rb") as f: 94 annotations = plistlib.load(f) 95 96 rois = [] 97 for image in annotations["Images"]: 98 for roi in image["ROIs"]: 99 label_id = ROI_NAME_TO_LABEL_ID.get(roi["Name"].strip().lower()) 100 if label_id is None: 101 continue 102 points = np.array([ 103 [float(v) for v in point.strip("()").split(",")] for point in roi["Point_px"] 104 ]) 105 rois.append((label_id, points)) 106 return rois 107 108 109def _rasterize_rois(rois, shape): 110 from skimage.draw import polygon 111 112 labels = np.zeros(shape, dtype="uint8") 113 for label_id, points in rois: 114 x, y = points[:, 0], points[:, 1] 115 if len(points) >= 3: 116 rr, cc = polygon(y, x, shape=shape) 117 else: 118 rr, cc = np.round(y).astype(int), np.round(x).astype(int) 119 valid = (rr >= 0) & (rr < shape[0]) & (cc >= 0) & (cc < shape[1]) 120 rr, cc = rr[valid], cc[valid] 121 labels[rr, cc] = label_id 122 return labels 123 124 125def _preprocess_inbreast(release_dir, preprocessed_dir): 126 import h5py 127 import pydicom 128 129 os.makedirs(preprocessed_dir, exist_ok=True) 130 131 dcm_paths = natsorted(glob(os.path.join(release_dir, "AllDICOMs", "*.dcm"))) 132 for dcm_path in tqdm(dcm_paths, desc="Preprocess INbreast"): 133 fname = os.path.basename(dcm_path) 134 case_id = fname.split("_")[0] 135 136 out_path = os.path.join(preprocessed_dir, f"{case_id}.h5") 137 if os.path.exists(out_path): 138 continue 139 140 raw = pydicom.dcmread(dcm_path).pixel_array 141 142 xml_path = os.path.join(release_dir, "AllXML", f"{case_id}.xml") 143 if os.path.exists(xml_path): 144 labels = _rasterize_rois(_parse_xml_rois(xml_path), raw.shape) 145 else: # A few images have no annotated lesions. 146 labels = np.zeros(raw.shape, dtype="uint8") 147 148 tmp_path = f"{out_path}.tmp" 149 with h5py.File(tmp_path, "w") as f: 150 f.create_dataset("raw", data=raw, compression="gzip") 151 f.create_dataset("labels", data=labels, compression="gzip") 152 os.rename(tmp_path, out_path) 153 154 return preprocessed_dir 155 156 157def get_inbreast_data(path: Union[os.PathLike, str], download: bool = False) -> str: 158 """Download the INbreast dataset. 159 160 Args: 161 path: Filepath to a folder where the data is downloaded for further processing. 162 download: Whether to download the data if it is not present. 163 164 Returns: 165 Filepath where the preprocessed data is stored. 166 """ 167 preprocessed_dir = os.path.join(path, "preprocessed") 168 if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0: 169 return preprocessed_dir 170 171 os.makedirs(path, exist_ok=True) 172 173 release_dir = os.path.join(path, "INbreast Release 1.0") 174 if not os.path.exists(release_dir): 175 zip_path = os.path.join(path, "inbreast-dataset.zip") 176 util.download_source_kaggle(path=path, dataset_name="ramanathansp20/inbreast-dataset", download=download) 177 _unzip_inbreast(zip_path, path) 178 179 return _preprocess_inbreast(release_dir, preprocessed_dir) 180 181 182def get_inbreast_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 183 """Get paths to the INbreast data. 184 185 Args: 186 path: Filepath to a folder where the data is downloaded for further processing. 187 download: Whether to download the data if it is not present. 188 189 Returns: 190 List of filepaths for the image data. 191 List of filepaths for the label data. 192 """ 193 data_dir = get_inbreast_data(path, download) 194 volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5"))) 195 assert len(volume_paths) > 0, f"Could not find any preprocessed files in '{data_dir}'." 196 return volume_paths, volume_paths 197 198 199def get_inbreast_dataset( 200 path: Union[os.PathLike, str], 201 patch_shape: Tuple[int, int], 202 resize_inputs: bool = False, 203 download: bool = False, 204 **kwargs 205) -> Dataset: 206 """Get the INbreast dataset for lesion segmentation in mammograms. 207 208 Args: 209 path: Filepath to a folder where the data is downloaded for further processing. 210 patch_shape: The patch shape to use for training. 211 resize_inputs: Whether to resize the inputs to the expected patch shape. 212 download: Whether to download the data if it is not present. 213 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 214 215 Returns: 216 The segmentation dataset. 217 """ 218 raw_paths, label_paths = get_inbreast_paths(path, download) 219 220 if resize_inputs: 221 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 222 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 223 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 224 ) 225 226 return torch_em.default_segmentation_dataset( 227 raw_paths=raw_paths, 228 raw_key="raw", 229 label_paths=label_paths, 230 label_key="labels", 231 patch_shape=patch_shape, 232 is_seg_dataset=True, 233 ndim=2, 234 **kwargs 235 ) 236 237 238def get_inbreast_loader( 239 path: Union[os.PathLike, str], 240 batch_size: int, 241 patch_shape: Tuple[int, int], 242 resize_inputs: bool = False, 243 download: bool = False, 244 **kwargs 245) -> DataLoader: 246 """Get the INbreast dataloader for lesion segmentation in mammograms. 247 248 Args: 249 path: Filepath to a folder where the data is downloaded for further processing. 250 batch_size: The batch size for training. 251 patch_shape: The patch shape to use for training. 252 resize_inputs: Whether to resize the inputs to the expected patch shape. 253 download: Whether to download the data if it is not present. 254 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 255 256 Returns: 257 The DataLoader. 258 """ 259 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 260 dataset = get_inbreast_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 261 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
158def get_inbreast_data(path: Union[os.PathLike, str], download: bool = False) -> str: 159 """Download the INbreast dataset. 160 161 Args: 162 path: Filepath to a folder where the data is downloaded for further processing. 163 download: Whether to download the data if it is not present. 164 165 Returns: 166 Filepath where the preprocessed data is stored. 167 """ 168 preprocessed_dir = os.path.join(path, "preprocessed") 169 if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0: 170 return preprocessed_dir 171 172 os.makedirs(path, exist_ok=True) 173 174 release_dir = os.path.join(path, "INbreast Release 1.0") 175 if not os.path.exists(release_dir): 176 zip_path = os.path.join(path, "inbreast-dataset.zip") 177 util.download_source_kaggle(path=path, dataset_name="ramanathansp20/inbreast-dataset", download=download) 178 _unzip_inbreast(zip_path, path) 179 180 return _preprocess_inbreast(release_dir, preprocessed_dir)
Download the INbreast dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the preprocessed data is stored.
183def get_inbreast_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 184 """Get paths to the INbreast data. 185 186 Args: 187 path: Filepath to a folder where the data is downloaded for further processing. 188 download: Whether to download the data if it is not present. 189 190 Returns: 191 List of filepaths for the image data. 192 List of filepaths for the label data. 193 """ 194 data_dir = get_inbreast_data(path, download) 195 volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5"))) 196 assert len(volume_paths) > 0, f"Could not find any preprocessed files in '{data_dir}'." 197 return volume_paths, volume_paths
Get paths to the INbreast data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
200def get_inbreast_dataset( 201 path: Union[os.PathLike, str], 202 patch_shape: Tuple[int, int], 203 resize_inputs: bool = False, 204 download: bool = False, 205 **kwargs 206) -> Dataset: 207 """Get the INbreast dataset for lesion segmentation in mammograms. 208 209 Args: 210 path: Filepath to a folder where the data is downloaded for further processing. 211 patch_shape: The patch shape to use for training. 212 resize_inputs: Whether to resize the inputs to the expected patch shape. 213 download: Whether to download the data if it is not present. 214 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 215 216 Returns: 217 The segmentation dataset. 218 """ 219 raw_paths, label_paths = get_inbreast_paths(path, download) 220 221 if resize_inputs: 222 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 223 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 224 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 225 ) 226 227 return torch_em.default_segmentation_dataset( 228 raw_paths=raw_paths, 229 raw_key="raw", 230 label_paths=label_paths, 231 label_key="labels", 232 patch_shape=patch_shape, 233 is_seg_dataset=True, 234 ndim=2, 235 **kwargs 236 )
Get the INbreast dataset for lesion segmentation in mammograms.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
239def get_inbreast_loader( 240 path: Union[os.PathLike, str], 241 batch_size: int, 242 patch_shape: Tuple[int, int], 243 resize_inputs: bool = False, 244 download: bool = False, 245 **kwargs 246) -> DataLoader: 247 """Get the INbreast dataloader for lesion segmentation in mammograms. 248 249 Args: 250 path: Filepath to a folder where the data is downloaded for further processing. 251 batch_size: The batch size for training. 252 patch_shape: The patch shape to use for training. 253 resize_inputs: Whether to resize the inputs to the expected patch shape. 254 download: Whether to download the data if it is not present. 255 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 256 257 Returns: 258 The DataLoader. 259 """ 260 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 261 dataset = get_inbreast_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 262 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the INbreast dataloader for lesion segmentation in mammograms.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize the inputs to the expected patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.