torch_em.data.datasets.medical.inbreast

The INbreast dataset contains annotations for lesion segmentation in full-field digital mammograms.

INbreast was acquired at the Breast Research Group, INESC Porto / Hospital de Sao Joao (Portugal) with a full-field digital mammography (FFDM) system, which distinguishes it from CBIS-DDSM (a re-scan of older film mammograms). It comprises 410 images from 115 cases (90 cases with both breasts, 4 images each; 25 mastectomy cases, 2 images each), each distributed as a DICOM file. Specialists outlined masses, calcifications, calcification clusters, spiculated regions, asymmetries and architectural distortions with the OsiriX viewer, and the resulting contours are exported per case as an OsiriX property list ('.xml', parseable with 'plistlib'). Each ROI lists its pixel coordinates under 'Point_px'; ROIs with 3 or more points are closed polygons (rasterized here with 'skimage.draw.polygon'), ROIs with 1 or 2 points are isolated pixel-level markers (usually individual, non-clustered calcifications).

The label ids used here are (see LABEL_IDS): 0 = background, 1 = mass, 2 = calcification, 3 = cluster (of calcifications), 4 = spiculated region, 5 = asymmetry, 6 = distortion. The 'Name' field of a few ROIs in the original xml files is misspelled or inconsistently capitalized (e.g. 'Assymetry', 'Espiculated Region', 'Calcifications'); these are normalized to the canonical names above. A handful of ROIs with an empty or placeholder name (e.g. 'Unnamed', 'Point 1') carry no lesion information and are skipped.

NOTE: The original INbreast distribution required a request form to the Breast Research Group and is not publicly downloadable anymore. This module downloads the 'INbreast Release 1.0' mirror hosted on Kaggle (https://www.kaggle.com/datasets/ramanathansp20/inbreast-dataset), which reproduces the original release folder structure ('AllDICOMs', 'AllXML', 'AllROI', 'MedicalReports', 'INbreast.xls', 'README.txt') including the per-lesion xml contours. Other INbreast mirrors on Kaggle (e.g. 'martholi/inbreast') only redistribute the DICOM images without the xml annotations and are not suitable for segmentation.

NOTE: Extracting the mirrored zip file with python's 'zipfile' can fail with a 'bad zipfile offset' / 'bad magic number' error, because the archive that Kaggle serves for this dataset has extra bytes prepended to it. This module extracts it with the 'unzip' command line tool instead, which recovers from this offset issue. Please make sure 'unzip' is available (it ships with most Linux distributions).

This dataset is from the publication https://doi.org/10.1016/j.acra.2011.09.014. Please cite it if you use this dataset in your research.

NOTE: The DICOM loading requires the 'pydicom' python package.

  1"""The INbreast dataset contains annotations for lesion segmentation in full-field digital mammograms.
  2
  3INbreast was acquired at the Breast Research Group, INESC Porto / Hospital de Sao Joao (Portugal) with a
  4full-field digital mammography (FFDM) system, which distinguishes it from CBIS-DDSM (a re-scan of older
  5film mammograms). It comprises 410 images from 115 cases (90 cases with both breasts, 4 images each; 25
  6mastectomy cases, 2 images each), each distributed as a DICOM file. Specialists outlined masses, calcifications,
  7calcification clusters, spiculated regions, asymmetries and architectural distortions with the OsiriX viewer,
  8and the resulting contours are exported per case as an OsiriX property list ('.xml', parseable with 'plistlib').
  9Each ROI lists its pixel coordinates under 'Point_px'; ROIs with 3 or more points are closed polygons
 10(rasterized here with 'skimage.draw.polygon'), ROIs with 1 or 2 points are isolated pixel-level markers
 11(usually individual, non-clustered calcifications).
 12
 13The label ids used here are (see `LABEL_IDS`): 0 = background, 1 = mass, 2 = calcification, 3 = cluster
 14(of calcifications), 4 = spiculated region, 5 = asymmetry, 6 = distortion. The 'Name' field of a few ROIs
 15in the original xml files is misspelled or inconsistently capitalized (e.g. 'Assymetry', 'Espiculated Region',
 16'Calcifications'); these are normalized to the canonical names above. A handful of ROIs with an empty or
 17placeholder name (e.g. 'Unnamed', 'Point 1') carry no lesion information and are skipped.
 18
 19NOTE: The original INbreast distribution required a request form to the Breast Research Group and is not
 20publicly downloadable anymore. This module downloads the 'INbreast Release 1.0' mirror hosted on Kaggle
 21(https://www.kaggle.com/datasets/ramanathansp20/inbreast-dataset), which reproduces the original release
 22folder structure ('AllDICOMs', 'AllXML', 'AllROI', 'MedicalReports', 'INbreast.xls', 'README.txt') including
 23the per-lesion xml contours. Other INbreast mirrors on Kaggle (e.g. 'martholi/inbreast') only redistribute
 24the DICOM images without the xml annotations and are not suitable for segmentation.
 25
 26NOTE: Extracting the mirrored zip file with python's 'zipfile' can fail with a 'bad zipfile offset' /
 27'bad magic number' error, because the archive that Kaggle serves for this dataset has extra bytes prepended
 28to it. This module extracts it with the 'unzip' command line tool instead, which recovers from this offset
 29issue. Please make sure 'unzip' is available (it ships with most Linux distributions).
 30
 31This dataset is from the publication https://doi.org/10.1016/j.acra.2011.09.014. Please cite it if you use
 32this dataset in your research.
 33
 34NOTE: The DICOM loading requires the 'pydicom' python package.
 35"""
 36
 37import os
 38from glob import glob
 39from shutil import which
 40from subprocess import run
 41from tqdm import tqdm
 42from natsort import natsorted
 43from typing import Union, Tuple, List
 44
 45import numpy as np
 46
 47from torch.utils.data import Dataset, DataLoader
 48
 49import torch_em
 50
 51from .. import util
 52
 53
 54LABEL_IDS = {
 55    "background": 0,
 56    "mass": 1,
 57    "calcification": 2,
 58    "cluster": 3,
 59    "spiculated_region": 4,
 60    "asymmetry": 5,
 61    "distortion": 6,
 62}
 63
 64# The 'Name' field of the xml ROIs is not fully standardized (typos, capitalization). This maps every
 65# variant that was found in the release to a canonical label id. Names that are not listed here (e.g. the
 66# empty string, 'Unnamed', 'Point 1') do not carry lesion information and are skipped.
 67ROI_NAME_TO_LABEL_ID = {
 68    "mass": LABEL_IDS["mass"],
 69    "calcification": LABEL_IDS["calcification"],
 70    "calcifications": LABEL_IDS["calcification"],
 71    "cluster": LABEL_IDS["cluster"],
 72    "spiculated region": LABEL_IDS["spiculated_region"],
 73    "espiculated region": LABEL_IDS["spiculated_region"],
 74    "asymmetry": LABEL_IDS["asymmetry"],
 75    "assymetry": LABEL_IDS["asymmetry"],
 76    "distortion": LABEL_IDS["distortion"],
 77}
 78
 79
 80def _unzip_inbreast(zip_path, dst):
 81    # 'zipfile' fails on this archive with a 'bad zipfile offset' error, because Kaggle serves it with
 82    # extra bytes prepended. The 'unzip' CLI recovers from this offset issue, so it is used instead of
 83    # 'torch_em.data.datasets.util.unzip'.
 84    if which("unzip") is None:
 85        raise RuntimeError("Need the 'unzip' CLI to extract the INbreast archive.")
 86    run(["unzip", "-q", "-o", zip_path, "-d", dst])
 87    os.remove(zip_path)
 88
 89
 90def _parse_xml_rois(xml_path):
 91    import plistlib
 92
 93    with open(xml_path, "rb") as f:
 94        annotations = plistlib.load(f)
 95
 96    rois = []
 97    for image in annotations["Images"]:
 98        for roi in image["ROIs"]:
 99            label_id = ROI_NAME_TO_LABEL_ID.get(roi["Name"].strip().lower())
100            if label_id is None:
101                continue
102            points = np.array([
103                [float(v) for v in point.strip("()").split(",")] for point in roi["Point_px"]
104            ])
105            rois.append((label_id, points))
106    return rois
107
108
109def _rasterize_rois(rois, shape):
110    from skimage.draw import polygon
111
112    labels = np.zeros(shape, dtype="uint8")
113    for label_id, points in rois:
114        x, y = points[:, 0], points[:, 1]
115        if len(points) >= 3:
116            rr, cc = polygon(y, x, shape=shape)
117        else:
118            rr, cc = np.round(y).astype(int), np.round(x).astype(int)
119            valid = (rr >= 0) & (rr < shape[0]) & (cc >= 0) & (cc < shape[1])
120            rr, cc = rr[valid], cc[valid]
121        labels[rr, cc] = label_id
122    return labels
123
124
125def _preprocess_inbreast(release_dir, preprocessed_dir):
126    import h5py
127    import pydicom
128
129    os.makedirs(preprocessed_dir, exist_ok=True)
130
131    dcm_paths = natsorted(glob(os.path.join(release_dir, "AllDICOMs", "*.dcm")))
132    for dcm_path in tqdm(dcm_paths, desc="Preprocess INbreast"):
133        fname = os.path.basename(dcm_path)
134        case_id = fname.split("_")[0]
135
136        out_path = os.path.join(preprocessed_dir, f"{case_id}.h5")
137        if os.path.exists(out_path):
138            continue
139
140        raw = pydicom.dcmread(dcm_path).pixel_array
141
142        xml_path = os.path.join(release_dir, "AllXML", f"{case_id}.xml")
143        if os.path.exists(xml_path):
144            labels = _rasterize_rois(_parse_xml_rois(xml_path), raw.shape)
145        else:  # A few images have no annotated lesions.
146            labels = np.zeros(raw.shape, dtype="uint8")
147
148        tmp_path = f"{out_path}.tmp"
149        with h5py.File(tmp_path, "w") as f:
150            f.create_dataset("raw", data=raw, compression="gzip")
151            f.create_dataset("labels", data=labels, compression="gzip")
152        os.rename(tmp_path, out_path)
153
154    return preprocessed_dir
155
156
157def get_inbreast_data(path: Union[os.PathLike, str], download: bool = False) -> str:
158    """Download the INbreast dataset.
159
160    Args:
161        path: Filepath to a folder where the data is downloaded for further processing.
162        download: Whether to download the data if it is not present.
163
164    Returns:
165        Filepath where the preprocessed data is stored.
166    """
167    preprocessed_dir = os.path.join(path, "preprocessed")
168    if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0:
169        return preprocessed_dir
170
171    os.makedirs(path, exist_ok=True)
172
173    release_dir = os.path.join(path, "INbreast Release 1.0")
174    if not os.path.exists(release_dir):
175        zip_path = os.path.join(path, "inbreast-dataset.zip")
176        util.download_source_kaggle(path=path, dataset_name="ramanathansp20/inbreast-dataset", download=download)
177        _unzip_inbreast(zip_path, path)
178
179    return _preprocess_inbreast(release_dir, preprocessed_dir)
180
181
182def get_inbreast_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
183    """Get paths to the INbreast data.
184
185    Args:
186        path: Filepath to a folder where the data is downloaded for further processing.
187        download: Whether to download the data if it is not present.
188
189    Returns:
190        List of filepaths for the image data.
191        List of filepaths for the label data.
192    """
193    data_dir = get_inbreast_data(path, download)
194    volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5")))
195    assert len(volume_paths) > 0, f"Could not find any preprocessed files in '{data_dir}'."
196    return volume_paths, volume_paths
197
198
199def get_inbreast_dataset(
200    path: Union[os.PathLike, str],
201    patch_shape: Tuple[int, int],
202    resize_inputs: bool = False,
203    download: bool = False,
204    **kwargs
205) -> Dataset:
206    """Get the INbreast dataset for lesion segmentation in mammograms.
207
208    Args:
209        path: Filepath to a folder where the data is downloaded for further processing.
210        patch_shape: The patch shape to use for training.
211        resize_inputs: Whether to resize the inputs to the expected patch shape.
212        download: Whether to download the data if it is not present.
213        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
214
215    Returns:
216        The segmentation dataset.
217    """
218    raw_paths, label_paths = get_inbreast_paths(path, download)
219
220    if resize_inputs:
221        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
222        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
223            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
224        )
225
226    return torch_em.default_segmentation_dataset(
227        raw_paths=raw_paths,
228        raw_key="raw",
229        label_paths=label_paths,
230        label_key="labels",
231        patch_shape=patch_shape,
232        is_seg_dataset=True,
233        ndim=2,
234        **kwargs
235    )
236
237
238def get_inbreast_loader(
239    path: Union[os.PathLike, str],
240    batch_size: int,
241    patch_shape: Tuple[int, int],
242    resize_inputs: bool = False,
243    download: bool = False,
244    **kwargs
245) -> DataLoader:
246    """Get the INbreast dataloader for lesion segmentation in mammograms.
247
248    Args:
249        path: Filepath to a folder where the data is downloaded for further processing.
250        batch_size: The batch size for training.
251        patch_shape: The patch shape to use for training.
252        resize_inputs: Whether to resize the inputs to the expected patch shape.
253        download: Whether to download the data if it is not present.
254        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
255
256    Returns:
257        The DataLoader.
258    """
259    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
260    dataset = get_inbreast_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
261    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
LABEL_IDS = {'background': 0, 'mass': 1, 'calcification': 2, 'cluster': 3, 'spiculated_region': 4, 'asymmetry': 5, 'distortion': 6}
ROI_NAME_TO_LABEL_ID = {'mass': 1, 'calcification': 2, 'calcifications': 2, 'cluster': 3, 'spiculated region': 4, 'espiculated region': 4, 'asymmetry': 5, 'assymetry': 5, 'distortion': 6}
def get_inbreast_data(path: Union[os.PathLike, str], download: bool = False) -> str:
158def get_inbreast_data(path: Union[os.PathLike, str], download: bool = False) -> str:
159    """Download the INbreast dataset.
160
161    Args:
162        path: Filepath to a folder where the data is downloaded for further processing.
163        download: Whether to download the data if it is not present.
164
165    Returns:
166        Filepath where the preprocessed data is stored.
167    """
168    preprocessed_dir = os.path.join(path, "preprocessed")
169    if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0:
170        return preprocessed_dir
171
172    os.makedirs(path, exist_ok=True)
173
174    release_dir = os.path.join(path, "INbreast Release 1.0")
175    if not os.path.exists(release_dir):
176        zip_path = os.path.join(path, "inbreast-dataset.zip")
177        util.download_source_kaggle(path=path, dataset_name="ramanathansp20/inbreast-dataset", download=download)
178        _unzip_inbreast(zip_path, path)
179
180    return _preprocess_inbreast(release_dir, preprocessed_dir)

Download the INbreast dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the preprocessed data is stored.

def get_inbreast_paths( path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
183def get_inbreast_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
184    """Get paths to the INbreast data.
185
186    Args:
187        path: Filepath to a folder where the data is downloaded for further processing.
188        download: Whether to download the data if it is not present.
189
190    Returns:
191        List of filepaths for the image data.
192        List of filepaths for the label data.
193    """
194    data_dir = get_inbreast_data(path, download)
195    volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5")))
196    assert len(volume_paths) > 0, f"Could not find any preprocessed files in '{data_dir}'."
197    return volume_paths, volume_paths

Get paths to the INbreast data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_inbreast_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
200def get_inbreast_dataset(
201    path: Union[os.PathLike, str],
202    patch_shape: Tuple[int, int],
203    resize_inputs: bool = False,
204    download: bool = False,
205    **kwargs
206) -> Dataset:
207    """Get the INbreast dataset for lesion segmentation in mammograms.
208
209    Args:
210        path: Filepath to a folder where the data is downloaded for further processing.
211        patch_shape: The patch shape to use for training.
212        resize_inputs: Whether to resize the inputs to the expected patch shape.
213        download: Whether to download the data if it is not present.
214        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
215
216    Returns:
217        The segmentation dataset.
218    """
219    raw_paths, label_paths = get_inbreast_paths(path, download)
220
221    if resize_inputs:
222        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
223        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
224            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
225        )
226
227    return torch_em.default_segmentation_dataset(
228        raw_paths=raw_paths,
229        raw_key="raw",
230        label_paths=label_paths,
231        label_key="labels",
232        patch_shape=patch_shape,
233        is_seg_dataset=True,
234        ndim=2,
235        **kwargs
236    )

Get the INbreast dataset for lesion segmentation in mammograms.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize the inputs to the expected patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_inbreast_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
239def get_inbreast_loader(
240    path: Union[os.PathLike, str],
241    batch_size: int,
242    patch_shape: Tuple[int, int],
243    resize_inputs: bool = False,
244    download: bool = False,
245    **kwargs
246) -> DataLoader:
247    """Get the INbreast dataloader for lesion segmentation in mammograms.
248
249    Args:
250        path: Filepath to a folder where the data is downloaded for further processing.
251        batch_size: The batch size for training.
252        patch_shape: The patch shape to use for training.
253        resize_inputs: Whether to resize the inputs to the expected patch shape.
254        download: Whether to download the data if it is not present.
255        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
256
257    Returns:
258        The DataLoader.
259    """
260    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
261    dataset = get_inbreast_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
262    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the INbreast dataloader for lesion segmentation in mammograms.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize the inputs to the expected patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.