torch_em.data.datasets.medical.acrin_6698

The ACRIN-6698 dataset contains annotations for functional tumor volume segmentation in dynamic contrast-enhanced (DCE) breast MRI, acquired as part of the ACRIN 6698 / I-SPY2 TRIAL.

The collection ships 1110 "VOLSER" (volume of signal enhancement ratio) DICOM-SEG objects: a binary mask of the functional tumor volume, obtained by thresholding a per-voxel signal enhancement ratio (SER) map computed from the pre- and post-contrast DCE-MRI phases. The mask is stored either "uni-lateral cropped" (covering one breast) or "bi-lateral" (covering both), cropped in-plane to the analyzed region. The collection also ships a second, unrelated family of 1103 "DWI SEG" objects (whole-tumor masks drawn on diffusion-weighted imaging derivative series); those are covered separately by medical.acrin_6698_dwi.

Unlike most DICOM-SEG collections handled elsewhere in this package, these SEG objects carry neither a ReferencedSeriesSequence nor a per-frame DerivationImageSequence: only a StudyInstanceUID and a FrameOfReferenceUID. Matching a mask to its source series therefore requires querying the TCIA REST API for all series of the mask's study and picking the DCE-MRI series with matching geometry, rather than the usual reference-UID lookup used in eg. nsclc_radiogenomics.py or adrenal_acc.py.

Inspecting real downloaded data (a random sample of 100 of the 1110 VOLSER masks, across both uni-lateral and bi-lateral cases) shows that every study contains exactly one MR series whose SeriesDescription is the mask's own description with "Analysis Mask" replaced by "SER" (eg. "ISPY2: VOLSER: uni-lateral cropped: Analysis Mask" pairs with "ISPY2: VOLSER: uni-lateral cropped: SER"). This SER series is the actual signal enhancement ratio map the mask was thresholded from: it shares the mask's FrameOfReferenceUID, in-plane PixelSpacing, ImageOrientationPatient and - since it is cropped to the same region - the exact same set of per-slice ImagePositionPatient values, unlike the full-field-of-view source DCE-MRI acquisition (eg. "ISPY2: Ax Vibrant PRE/POST"), which shares the frame of reference and in-plane geometry but not the cropped extent. This module therefore uses the description-matched SER series as the raw image paired with each mask. As a safety net (in case the description-based match is ever wrong or absent for some study), the matched candidate's geometry is verified by resampling the mask onto the candidate's voxel grid (see _resample_labels in adrenal_acc.py) and checking that (nearly) all mask foreground voxels land inside it; if this fails, every other MR series of the study with the same FrameOfReferenceUID is tried in turn, and the case is skipped (with a warning) if none of them pass the geometry check either. In the validated sample of 100 masks, the description-based match succeeded and passed geometry verification in all 100 cases (0% failure rate); a larger production run may still see occasional skips for studies with unusual or missing SER series, which are handled gracefully rather than raised as errors.

Two further data quirks were found by inspecting the actual downloaded pixel data of 5 validated cases and had to be handled explicitly:

  • The SEG objects use SegmentationType = FRACTIONAL with SegmentationFractionalType = OCCUPANCY, not BINARY, even though SegmentSequence lists only a single segment, and its Content Description reads "Bit-map mask of SER tumor segmentation". This is NOT a 0-255 fractional occupancy value: it is an inverse, bit-encoded exclusion mask, as documented in the "Analysis masks from FTV processing" support document linked from the TCIA collection page (https://www.cancerimagingarchive.net/collection/acrin-6698/, filename 'Analysis-mask-files-description.v20211020.docx', by the I-SPY imaging group, UCSF - fetched and parsed directly, not guessed). Per that document: "The masks are INVERSE masks, in that a mask value of 0 indicates that a voxel was included in the measured functional tumor volume (FTV)". Each of up to 5 masking/exclusion steps that removed a voxel from the FTV sets one bit of the byte value: 1 = failed the PE (percent enhancement) threshold, 2 = failed the 3D minimum-neighbor-count (MNC) connectivity filter, 32 = outside the manually-drawn rectangular VOI, 64 = manually OMITted by a trained observer, and a background-intensity-threshold exclusion bit (documented as 8, but empirically 16 in the ACRIN-6698 SEG files actually inspected - see below). A voxel excluded by several steps has the bitwise OR of their values; only {0, 1, 2, 16, 17, 32, 33, 34, 48, 49, ...}-style combinations are therefore valid, and this matches byte-for-byte what is observed in the 5 validated cases (exactly {0, 1, 2, 17, 32, 33, 34, 49}, ie. every combination of {PE, background, VOI} that occurred in the sample, decomposing cleanly as 17 = 16 (background) + 1 (PE), 33 = 32 (VOI) + 1 (PE), 34 = 32 (VOI) + 2 (MNC), 49 = 32 + 16 + 1). This also explains the sharp-edged rectangular sub-regions seen in an earlier (incorrect) version of this module that treated every non-background-mode voxel as foreground: those rectangles are the boundary of the manually-drawn VOI itself (voxels just inside vs. just outside it differ only in whether the 32 bit is set), not a real anatomical structure. This module now instead uses the documented semantics directly: foreground (the FTV) is exactly the voxels with value 0. Verified for the 5 validated cases: the resulting mask is a small, spatially compact, contiguous 3D blob (or, in 1/5 cases, empty - a legitimate "no residual enhancing tumor" result, since ACRIN-6698 includes post-treatment timepoints) whose per-slice area traces a smooth, single-peaked profile, and which spatially coincides exactly with the locations of elevated signal in the matched SER map (see the second quirk below and check_acrin_6698.py) - unlike the rectangle-contaminated mask from treating "not background" as foreground.
  • The matched SER (and other VOLSER-derived) MR series store their quantitative ratio values with a very small RescaleSlope (0.001 in the validated cases, ie. the raw stored int16 pixel value is the SER ratio times 1000). Applying the slope/intercept and rounding to an integer volume, as adrenal_acc.py does for CT Hounsfield units, collapses nearly the entire dynamic range to {0, 1, 2, 3} and destroys the image. This module therefore keeps the raw stored pixel values as-is for the source volume, without applying the rescale slope/intercept. Separately, and not a bug: the SER maps themselves are extremely sparse (only 0.02-1.9% of voxels nonzero in the validated cases) - by design, per the same support document, the background/non-enhancing/non-tissue voxels are zeroed out by the upstream processing pipeline before this derived series is exported, so most of the cropped volume being exactly 0 is a genuine property of the data, not a display, windowing or download issue.

Because the collection is large (1110 relevant SEG series across 385 patients), use max_cases or case_ids to only download and preprocess a subset.

NOTE: This requires the pydicom python package.

The dataset is located at https://www.cancerimagingarchive.net/collection/acrin-6698/ and is distributed under the CC BY 4.0 license.

This dataset is from the publication https://doi.org/10.1148/radiol.2018180273. The data was released at https://doi.org/10.7937/tcia.kk02-6d95. Please cite it if you use this dataset in your research.

  1"""The ACRIN-6698 dataset contains annotations for functional tumor volume segmentation in dynamic
  2contrast-enhanced (DCE) breast MRI, acquired as part of the ACRIN 6698 / I-SPY2 TRIAL.
  3
  4The collection ships 1110 "VOLSER" (volume of signal enhancement ratio) DICOM-SEG objects: a binary mask
  5of the functional tumor volume, obtained by thresholding a per-voxel signal enhancement ratio (SER) map
  6computed from the pre- and post-contrast DCE-MRI phases. The mask is stored either "uni-lateral cropped"
  7(covering one breast) or "bi-lateral" (covering both), cropped in-plane to the analyzed region. The
  8collection also ships a second, unrelated family of 1103 "DWI SEG" objects (whole-tumor masks drawn on
  9diffusion-weighted imaging derivative series); those are covered separately by `medical.acrin_6698_dwi`.
 10
 11Unlike most DICOM-SEG collections handled elsewhere in this package, these SEG objects carry neither a
 12`ReferencedSeriesSequence` nor a per-frame `DerivationImageSequence`: only a `StudyInstanceUID` and a
 13`FrameOfReferenceUID`. Matching a mask to its source series therefore requires querying the TCIA REST API
 14for all series of the mask's study and picking the DCE-MRI series with matching geometry, rather than the
 15usual reference-UID lookup used in eg. `nsclc_radiogenomics.py` or `adrenal_acc.py`.
 16
 17Inspecting real downloaded data (a random sample of 100 of the 1110 VOLSER masks, across both uni-lateral
 18and bi-lateral cases) shows that every study contains exactly one MR series whose `SeriesDescription` is the
 19mask's own description with "Analysis Mask" replaced by "SER" (eg. "ISPY2: VOLSER: uni-lateral cropped:
 20Analysis Mask" pairs with "ISPY2: VOLSER: uni-lateral cropped: SER"). This SER series is the actual signal
 21enhancement ratio map the mask was thresholded from: it shares the mask's `FrameOfReferenceUID`, in-plane
 22`PixelSpacing`, `ImageOrientationPatient` and - since it is cropped to the same region - the exact same set
 23of per-slice `ImagePositionPatient` values, unlike the full-field-of-view source DCE-MRI acquisition (eg.
 24"ISPY2: Ax Vibrant PRE/POST"), which shares the frame of reference and in-plane geometry but not the
 25cropped extent. This module therefore uses the description-matched SER series as the raw image paired with
 26each mask. As a safety net (in case the description-based match is ever wrong or absent for some study),
 27the matched candidate's geometry is verified by resampling the mask onto the candidate's voxel grid (see
 28`_resample_labels` in `adrenal_acc.py`) and checking that (nearly) all mask foreground voxels land inside
 29it; if this fails, every other MR series of the study with the same `FrameOfReferenceUID` is tried in turn,
 30and the case is skipped (with a warning) if none of them pass the geometry check either. In the validated
 31sample of 100 masks, the description-based match succeeded and passed geometry verification in all 100
 32cases (0% failure rate); a larger production run may still see occasional skips for studies with unusual or
 33missing SER series, which are handled gracefully rather than raised as errors.
 34
 35Two further data quirks were found by inspecting the actual downloaded pixel data of 5 validated cases and
 36had to be handled explicitly:
 37- The SEG objects use `SegmentationType = FRACTIONAL` with `SegmentationFractionalType = OCCUPANCY`, not
 38  `BINARY`, even though `SegmentSequence` lists only a single segment, and its `Content Description` reads
 39  "Bit-map mask of SER tumor segmentation". This is NOT a 0-255 fractional occupancy value: it is an
 40  *inverse*, bit-encoded exclusion mask, as documented in the "Analysis masks from FTV processing" support
 41  document linked from the TCIA collection page (https://www.cancerimagingarchive.net/collection/acrin-6698/,
 42  filename 'Analysis-mask-files-description.v20211020.docx', by the I-SPY imaging group, UCSF - fetched and
 43  parsed directly, not guessed). Per that document: "The masks are INVERSE masks, in that a mask value of 0
 44  indicates that a voxel was included in the measured functional tumor volume (FTV)". Each of up to 5
 45  masking/exclusion steps that removed a voxel from the FTV sets one bit of the byte value: 1 = failed the
 46  PE (percent enhancement) threshold, 2 = failed the 3D minimum-neighbor-count (MNC) connectivity filter,
 47  32 = outside the manually-drawn rectangular VOI, 64 = manually OMITted by a trained observer, and a
 48  background-intensity-threshold exclusion bit (documented as 8, but empirically 16 in the ACRIN-6698 SEG
 49  files actually inspected - see below). A voxel excluded by several steps has the bitwise OR of their
 50  values; only `{0, 1, 2, 16, 17, 32, 33, 34, 48, 49, ...}`-style combinations are therefore valid, and this
 51  matches byte-for-byte what is observed in the 5 validated cases (exactly `{0, 1, 2, 17, 32, 33, 34, 49}`,
 52  ie. every combination of {PE, background, VOI} that occurred in the sample, decomposing cleanly as
 53  17 = 16 (background) + 1 (PE), 33 = 32 (VOI) + 1 (PE), 34 = 32 (VOI) + 2 (MNC), 49 = 32 + 16 + 1). This
 54  also explains the sharp-edged rectangular sub-regions seen in an earlier (incorrect) version of this
 55  module that treated every non-background-mode voxel as foreground: those rectangles are the boundary of
 56  the manually-drawn VOI itself (voxels just inside vs. just outside it differ only in whether the 32 bit is
 57  set), not a real anatomical structure. This module now instead uses the documented semantics directly:
 58  foreground (the FTV) is exactly the voxels with value 0. Verified for the 5 validated cases: the resulting
 59  mask is a small, spatially compact, contiguous 3D blob (or, in 1/5 cases, empty - a legitimate "no residual
 60  enhancing tumor" result, since ACRIN-6698 includes post-treatment timepoints) whose per-slice area traces
 61  a smooth, single-peaked profile, and which spatially coincides exactly with the locations of elevated
 62  signal in the matched SER map (see the second quirk below and `check_acrin_6698.py`) - unlike the
 63  rectangle-contaminated mask from treating "not background" as foreground.
 64- The matched SER (and other VOLSER-derived) MR series store their quantitative ratio values with a very
 65  small `RescaleSlope` (0.001 in the validated cases, ie. the raw stored int16 pixel value is the SER ratio
 66  times 1000). Applying the slope/intercept and rounding to an integer volume, as `adrenal_acc.py` does for
 67  CT Hounsfield units, collapses nearly the entire dynamic range to {0, 1, 2, 3} and destroys the image. This
 68  module therefore keeps the raw stored pixel values as-is for the source volume, without applying the
 69  rescale slope/intercept. Separately, and not a bug: the SER maps themselves are extremely sparse (only
 70  0.02-1.9% of voxels nonzero in the validated cases) - by design, per the same support document, the
 71  background/non-enhancing/non-tissue voxels are zeroed out by the upstream processing pipeline before this
 72  derived series is exported, so most of the cropped volume being exactly 0 is a genuine property of the
 73  data, not a display, windowing or download issue.
 74
 75Because the collection is large (1110 relevant SEG series across 385 patients), use `max_cases` or
 76`case_ids` to only download and preprocess a subset.
 77
 78NOTE: This requires the pydicom python package.
 79
 80The dataset is located at https://www.cancerimagingarchive.net/collection/acrin-6698/ and is distributed
 81under the CC BY 4.0 license.
 82
 83This dataset is from the publication https://doi.org/10.1148/radiol.2018180273.
 84The data was released at https://doi.org/10.7937/tcia.kk02-6d95.
 85Please cite it if you use this dataset in your research.
 86"""
 87
 88import os
 89from glob import glob
 90from tqdm import tqdm
 91from warnings import warn
 92from natsort import natsorted
 93from typing import Union, Tuple, List, Optional
 94
 95import numpy as np
 96import requests
 97
 98from torch.utils.data import Dataset, DataLoader
 99
100import torch_em
101
102from .adrenal_acc import _resample_labels
103from .. import util
104
105
106COLLECTION = "ACRIN-6698"
107
108LABEL_IDS = {"functional_tumor_volume": 1}
109
110# The minimum fraction of the mask's foreground voxels that must land inside a candidate source series
111# (after resampling onto its voxel grid) for that candidate to be accepted as the match.
112MIN_GEOMETRY_OVERLAP = 0.9
113
114
115def _load_dicom_volume(series_dir):
116    """Stack a DICOM series into a volume with axes (z, y, x) and slices sorted along the slice normal.
117
118    Unlike `adrenal_acc._load_dicom_volume`, the rescale slope/intercept is NOT applied here: see the
119    module docstring for why this matters for the ACRIN-6698 VOLSER-derived MR series (SER, PE2, PE6, ...).
120
121    Returns the volume and the affine matrix that maps voxel indices (z, y, x) to DICOM patient coordinates.
122    """
123    import pydicom
124
125    slices = [pydicom.dcmread(dcm_path) for dcm_path in natsorted(glob(os.path.join(series_dir, "*.dcm")))]
126    orientation = np.array([float(v) for v in slices[0].ImageOrientationPatient])
127    row_dir, col_dir = orientation[:3], orientation[3:]
128    normal = np.cross(row_dir, col_dir)
129    slices.sort(key=lambda dcm: np.dot([float(v) for v in dcm.ImagePositionPatient], normal))
130
131    volume = np.stack([dcm.pixel_array for dcm in slices])
132
133    positions = np.array([[float(v) for v in dcm.ImagePositionPatient] for dcm in slices])
134    spacing = [float(v) for v in slices[0].PixelSpacing]  # The spacing between rows and between columns.
135    affine = np.eye(4)
136    affine[:3, 0] = (positions[-1] - positions[0]) / (len(slices) - 1)
137    affine[:3, 1] = col_dir * spacing[0]
138    affine[:3, 2] = row_dir * spacing[1]
139    affine[:3, 3] = positions[0]
140    return volume, affine
141
142
143def _load_dicom_seg(seg_path):
144    """Load a VOLSER analysis mask DICOM-SEG object as a label volume with axes (z, y, x).
145
146    Unlike `adrenal_acc._load_dicom_seg`, foreground is not determined by a nonzero pixel value: the
147    ACRIN-6698 VOLSER masks are documented (see the module docstring) as INVERSE, bit-encoded exclusion
148    masks, where a voxel value of exactly 0 means it was included in the functional tumor volume (FTV) and
149    any nonzero value means it was excluded by one or more processing steps.
150
151    Returns the label volume and the affine matrix that maps its voxel indices to DICOM patient coordinates.
152    """
153    import pydicom
154
155    seg = pydicom.dcmread(seg_path)
156    frames = seg.pixel_array
157    if frames.ndim == 2:  # A segmentation with a single frame.
158        frames = frames[None]
159
160    shared_group = seg.SharedFunctionalGroupsSequence[0]
161    orientation = np.array([float(v) for v in shared_group.PlaneOrientationSequence[0].ImageOrientationPatient])
162    row_dir, col_dir = orientation[:3], orientation[3:]
163    normal = np.cross(row_dir, col_dir)
164    pixel_measures = shared_group.PixelMeasuresSequence[0]
165    spacing = [float(v) for v in pixel_measures.PixelSpacing]  # The spacing between rows and between columns.
166
167    frame_groups = seg.PerFrameFunctionalGroupsSequence
168    positions = np.array([[float(v) for v in g.PlanePositionSequence[0].ImagePositionPatient] for g in frame_groups])
169    projections = positions @ normal
170    if "SpacingBetweenSlices" in pixel_measures:
171        slice_spacing = float(pixel_measures.SpacingBetweenSlices)
172    elif len(projections) > 1:
173        slice_spacing = np.min(np.diff(np.unique(np.round(projections, 3))))
174    else:
175        slice_spacing = float(pixel_measures.SliceThickness)
176    slice_ids = np.round((projections - projections.min()) / slice_spacing).astype("int")
177
178    # A value of 0 means the voxel was included in the FTV, see the module docstring; any nonzero value
179    # means it was excluded (the bitwise OR of one or more masking-step codes).
180    segment_number = int(seg.SegmentSequence[0].SegmentNumber)
181
182    labels = np.zeros((slice_ids.max() + 1, seg.Rows, seg.Columns), dtype="uint8")
183    for frame, slice_id in zip(frames, slice_ids):
184        labels[slice_id][frame == 0] = segment_number
185
186    affine = np.eye(4)
187    affine[:3, 0] = normal * slice_spacing
188    affine[:3, 1] = col_dir * spacing[0]
189    affine[:3, 2] = row_dir * spacing[1]
190    affine[:3, 3] = positions[np.argmin(projections)]
191    return labels, affine
192
193
194def _get_seg_metadata(path: str, download: bool) -> List[dict]:
195    """Query the TCIA REST API for the metadata of all SEG series of the collection, without downloading images."""
196    import json
197
198    metadata_path = os.path.join(path, "acrin_6698_seg_series.json")
199    if not os.path.exists(metadata_path):
200        if not download:
201            raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
202        response = requests.get(util.NBIA_API_URL + "getSeries", params={"Collection": COLLECTION, "Modality": "SEG"})
203        response.raise_for_status()
204        with open(metadata_path, "w") as f:
205            json.dump(response.json(), f, indent=2)
206
207    with open(metadata_path, "r") as f:
208        return json.load(f)
209
210
211def _is_volser_mask(series_description: str) -> bool:
212    return "VOLSER" in series_description and "Analysis Mask" in series_description
213
214
215def _get_study_series(study_uid: str, cache: dict) -> List[dict]:
216    """Query the TCIA REST API for the metadata of all series of a study, cached per study."""
217    if study_uid not in cache:
218        response = requests.get(util.NBIA_API_URL + "getSeries", params={"StudyInstanceUID": study_uid})
219        response.raise_for_status()
220        cache[study_uid] = response.json()
221    return cache[study_uid]
222
223
224def _match_by_description(seg_description: str, study_series: List[dict]) -> Optional[str]:
225    """Find the SER series whose description matches the mask's description (see the module docstring)."""
226    prefix = seg_description.replace("Analysis Mask", "SER").strip()
227    matches = [
228        series["SeriesInstanceUID"] for series in study_series
229        if series["Modality"] == "MR" and series["SeriesDescription"].strip() == prefix
230    ]
231    return matches[0] if len(matches) == 1 else None
232
233
234def _candidate_mr_uids(seg_description: str, study_series: List[dict]) -> List[str]:
235    """List the candidate source series of a study, the description match (if any) tried first."""
236    primary = _match_by_description(seg_description, study_series)
237    others = [
238        series["SeriesInstanceUID"] for series in study_series
239        if series["Modality"] == "MR" and series["SeriesInstanceUID"] != primary
240    ]
241    return ([primary] if primary is not None else []) + others
242
243
244def _resolve_source_volume(seg_path, seg_labels, seg_affine, frame_of_reference, mr_uids, dicom_dir, download):
245    """Find the source MR series matching a mask's geometry and return its volume, aligned labels and affine.
246
247    Tries the candidate series in order (the description-matched one first, see `_candidate_mr_uids`),
248    downloading each on demand, and accepts the first one whose frame of reference matches and whose
249    resampled mask overlap exceeds `MIN_GEOMETRY_OVERLAP`. A mask can legitimately have no foreground voxels
250    at all (an empty functional tumor volume, eg. a complete response at a later trial timepoint, see the
251    module docstring), in which case the overlap fraction is undefined; the first candidate with a matching
252    frame of reference and pixel grid (`Rows`/`Columns`) is accepted instead, without downloading and trying
253    every other series of the study. Returns (None, None, None) if none match.
254    """
255    import pydicom
256
257    n_foreground = int((seg_labels > 0).sum())
258    seg_dcm = pydicom.dcmread(seg_path, stop_before_pixels=True)
259    seg_shape = (int(seg_dcm.Rows), int(seg_dcm.Columns))
260
261    for mr_uid in mr_uids:
262        mr_dir = os.path.join(dicom_dir, mr_uid)
263        if not glob(os.path.join(mr_dir, "*.dcm")):
264            if not download:
265                continue
266            util.download_tcia_series([mr_uid], dst=dicom_dir, csv_filename=os.path.join(dicom_dir, "..", "acrin_6698_mr_extra"))  # noqa
267        dcm_paths = glob(os.path.join(mr_dir, "*.dcm"))
268        if not dcm_paths:
269            continue
270
271        first_dcm = pydicom.dcmread(dcm_paths[0], stop_before_pixels=True)
272        if str(getattr(first_dcm, "FrameOfReferenceUID", None)) != frame_of_reference:
273            continue
274
275        if n_foreground == 0:
276            if (int(first_dcm.Rows), int(first_dcm.Columns)) != seg_shape:
277                continue
278            volume, mr_affine = _load_dicom_volume(mr_dir)
279            labels = np.zeros(volume.shape, dtype=seg_labels.dtype)
280            return volume, labels, mr_affine
281
282        volume, mr_affine = _load_dicom_volume(mr_dir)
283        labels = _resample_labels(seg_labels, seg_affine, volume.shape, mr_affine)
284        overlap = int((labels > 0).sum()) / n_foreground
285        if overlap >= MIN_GEOMETRY_OVERLAP:
286            return volume, labels, mr_affine
287
288    return None, None, None
289
290
291def _preprocess_acrin_6698(seg_series: List[dict], dicom_dir: str, preprocessed_dir: str, download: bool) -> None:
292    import h5py
293    import pydicom
294
295    os.makedirs(preprocessed_dir, exist_ok=True)
296    study_series_cache = {}
297    for series in tqdm(seg_series, desc="Preprocess ACRIN-6698"):
298        seg_uid = series["SeriesInstanceUID"]
299        out_path = os.path.join(preprocessed_dir, f"{seg_uid}.h5")
300        if os.path.exists(out_path):
301            continue
302
303        seg_paths = glob(os.path.join(dicom_dir, seg_uid, "*.dcm"))
304        if not seg_paths:
305            warn(f"Skipping {seg_uid}, whose SEG DICOM data could not be found.")
306            continue
307
308        seg_dcm = pydicom.dcmread(seg_paths[0], stop_before_pixels=True)
309        frame_of_reference = str(seg_dcm.FrameOfReferenceUID)
310
311        study_series = _get_study_series(str(seg_dcm.StudyInstanceUID), study_series_cache)
312        mr_uids = _candidate_mr_uids(series["SeriesDescription"], study_series)
313        if not mr_uids:
314            warn(f"Skipping {seg_uid}, whose study has no MR series to match against.")
315            continue
316
317        seg_labels, seg_affine = _load_dicom_seg(seg_paths[0])
318        volume, labels, _ = _resolve_source_volume(
319            seg_paths[0], seg_labels, seg_affine, frame_of_reference, mr_uids, dicom_dir, download
320        )
321        if volume is None:
322            warn(f"Skipping {seg_uid}, for which no MR series with matching geometry could be found.")
323            continue
324
325        labels = (labels > 0).astype("uint8") * LABEL_IDS["functional_tumor_volume"]
326        with h5py.File(out_path, "w") as f:
327            f.create_dataset("raw", data=volume, compression="gzip")
328            f.create_dataset("labels", data=labels, compression="gzip")
329
330
331def get_acrin_6698_data(
332    path: Union[os.PathLike, str],
333    max_cases: Optional[int] = None,
334    case_ids: Optional[List[str]] = None,
335    download: bool = False,
336) -> str:
337    """Download the ACRIN-6698 dataset.
338
339    Args:
340        path: Filepath to a folder where the data is downloaded for further processing.
341        max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic
342            order of their series instance UID. Only the SEG series and their matched source MR series are
343            downloaded. Mutually exclusive with `case_ids`.
344        case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the
345            'acrin_6698_seg_series.json' metadata file written to `path` for the available UIDs and their
346            descriptions). Mutually exclusive with `max_cases`.
347        download: Whether to download the data if it is not present.
348
349    Returns:
350        Filepath where the preprocessed data is stored.
351    """
352    assert max_cases is None or case_ids is None, "'max_cases' and 'case_ids' are mutually exclusive."
353
354    # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes.
355    preprocessed_dir = os.path.join(path, "preprocessed")
356    os.makedirs(path, exist_ok=True)
357
358    seg_series = _get_seg_metadata(path, download)
359    seg_series = [series for series in seg_series if _is_volser_mask(series["SeriesDescription"])]
360    seg_series = natsorted(seg_series, key=lambda series: series["SeriesInstanceUID"])
361
362    if max_cases is not None:
363        seg_series = seg_series[:max_cases]
364    elif case_ids is not None:
365        seg_series = [series for series in seg_series if series["SeriesInstanceUID"] in case_ids]
366
367    dicom_dir = os.path.join(path, "dicom")
368    if download:
369        seg_uids = [series["SeriesInstanceUID"] for series in seg_series]
370        util.download_tcia_series(seg_uids, dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_seg"))
371
372        # The primary (description-matched) candidate source series of each mask's study is downloaded here
373        # already, so that the preprocessing loop below only has to download extra candidates for the rare
374        # case where the primary candidate's geometry does not match (see the module docstring).
375        study_series_cache = {}
376        primary_mr_uids = set()
377        for series in seg_series:
378            seg_paths = glob(os.path.join(dicom_dir, series["SeriesInstanceUID"], "*.dcm"))
379            if not seg_paths:
380                continue
381            import pydicom
382            study_uid = str(pydicom.dcmread(seg_paths[0], stop_before_pixels=True).StudyInstanceUID)
383            study_series = _get_study_series(study_uid, study_series_cache)
384            primary_uid = _match_by_description(series["SeriesDescription"], study_series)
385            if primary_uid is not None:
386                primary_mr_uids.add(primary_uid)
387        util.download_tcia_series(
388            sorted(primary_mr_uids), dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_mr")
389        )
390    elif not glob(os.path.join(dicom_dir, "*", "*.dcm")):
391        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
392
393    _preprocess_acrin_6698(seg_series, dicom_dir, preprocessed_dir, download)
394    return preprocessed_dir
395
396
397def get_acrin_6698_paths(
398    path: Union[os.PathLike, str],
399    max_cases: Optional[int] = None,
400    case_ids: Optional[List[str]] = None,
401    download: bool = False,
402) -> List[str]:
403    """Get paths to the ACRIN-6698 data.
404
405    Args:
406        path: Filepath to a folder where the data is downloaded for further processing.
407        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
408        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
409        download: Whether to download the data if it is not present.
410
411    Returns:
412        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
413    """
414    preprocessed_dir = get_acrin_6698_data(path, max_cases, case_ids, download)
415    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
416    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'."
417    return volume_paths
418
419
420def get_acrin_6698_dataset(
421    path: Union[os.PathLike, str],
422    patch_shape: Tuple[int, ...],
423    resize_inputs: bool = False,
424    max_cases: Optional[int] = None,
425    case_ids: Optional[List[str]] = None,
426    download: bool = False,
427    **kwargs
428) -> Dataset:
429    """Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI.
430
431    Args:
432        path: Filepath to a folder where the data is downloaded for further processing.
433        patch_shape: The patch shape to use for training.
434        resize_inputs: Whether to resize inputs to the desired patch shape.
435        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
436        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
437        download: Whether to download the data if it is not present.
438        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
439
440    Returns:
441        The segmentation dataset.
442    """
443    volume_paths = get_acrin_6698_paths(path, max_cases, case_ids, download)
444
445    if resize_inputs:
446        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
447        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
448            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
449        )
450
451    return torch_em.default_segmentation_dataset(
452        raw_paths=volume_paths,
453        raw_key="raw",
454        label_paths=volume_paths,
455        label_key="labels",
456        patch_shape=patch_shape,
457        is_seg_dataset=True,
458        **kwargs
459    )
460
461
462def get_acrin_6698_loader(
463    path: Union[os.PathLike, str],
464    batch_size: int,
465    patch_shape: Tuple[int, ...],
466    resize_inputs: bool = False,
467    max_cases: Optional[int] = None,
468    case_ids: Optional[List[str]] = None,
469    download: bool = False,
470    **kwargs
471) -> DataLoader:
472    """Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI.
473
474    Args:
475        path: Filepath to a folder where the data is downloaded for further processing.
476        batch_size: The batch size for training.
477        patch_shape: The patch shape to use for training.
478        resize_inputs: Whether to resize inputs to the desired patch shape.
479        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
480        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
481        download: Whether to download the data if it is not present.
482        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
483
484    Returns:
485        The DataLoader.
486    """
487    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
488    dataset = get_acrin_6698_dataset(path, patch_shape, resize_inputs, max_cases, case_ids, download, **ds_kwargs)
489    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
COLLECTION = 'ACRIN-6698'
LABEL_IDS = {'functional_tumor_volume': 1}
MIN_GEOMETRY_OVERLAP = 0.9
def get_acrin_6698_data( path: Union[os.PathLike, str], max_cases: Optional[int] = None, case_ids: Optional[List[str]] = None, download: bool = False) -> str:
332def get_acrin_6698_data(
333    path: Union[os.PathLike, str],
334    max_cases: Optional[int] = None,
335    case_ids: Optional[List[str]] = None,
336    download: bool = False,
337) -> str:
338    """Download the ACRIN-6698 dataset.
339
340    Args:
341        path: Filepath to a folder where the data is downloaded for further processing.
342        max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic
343            order of their series instance UID. Only the SEG series and their matched source MR series are
344            downloaded. Mutually exclusive with `case_ids`.
345        case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the
346            'acrin_6698_seg_series.json' metadata file written to `path` for the available UIDs and their
347            descriptions). Mutually exclusive with `max_cases`.
348        download: Whether to download the data if it is not present.
349
350    Returns:
351        Filepath where the preprocessed data is stored.
352    """
353    assert max_cases is None or case_ids is None, "'max_cases' and 'case_ids' are mutually exclusive."
354
355    # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes.
356    preprocessed_dir = os.path.join(path, "preprocessed")
357    os.makedirs(path, exist_ok=True)
358
359    seg_series = _get_seg_metadata(path, download)
360    seg_series = [series for series in seg_series if _is_volser_mask(series["SeriesDescription"])]
361    seg_series = natsorted(seg_series, key=lambda series: series["SeriesInstanceUID"])
362
363    if max_cases is not None:
364        seg_series = seg_series[:max_cases]
365    elif case_ids is not None:
366        seg_series = [series for series in seg_series if series["SeriesInstanceUID"] in case_ids]
367
368    dicom_dir = os.path.join(path, "dicom")
369    if download:
370        seg_uids = [series["SeriesInstanceUID"] for series in seg_series]
371        util.download_tcia_series(seg_uids, dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_seg"))
372
373        # The primary (description-matched) candidate source series of each mask's study is downloaded here
374        # already, so that the preprocessing loop below only has to download extra candidates for the rare
375        # case where the primary candidate's geometry does not match (see the module docstring).
376        study_series_cache = {}
377        primary_mr_uids = set()
378        for series in seg_series:
379            seg_paths = glob(os.path.join(dicom_dir, series["SeriesInstanceUID"], "*.dcm"))
380            if not seg_paths:
381                continue
382            import pydicom
383            study_uid = str(pydicom.dcmread(seg_paths[0], stop_before_pixels=True).StudyInstanceUID)
384            study_series = _get_study_series(study_uid, study_series_cache)
385            primary_uid = _match_by_description(series["SeriesDescription"], study_series)
386            if primary_uid is not None:
387                primary_mr_uids.add(primary_uid)
388        util.download_tcia_series(
389            sorted(primary_mr_uids), dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_mr")
390        )
391    elif not glob(os.path.join(dicom_dir, "*", "*.dcm")):
392        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
393
394    _preprocess_acrin_6698(seg_series, dicom_dir, preprocessed_dir, download)
395    return preprocessed_dir

Download the ACRIN-6698 dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic order of their series instance UID. Only the SEG series and their matched source MR series are downloaded. Mutually exclusive with case_ids.
  • case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the 'acrin_6698_seg_series.json' metadata file written to path for the available UIDs and their descriptions). Mutually exclusive with max_cases.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the preprocessed data is stored.

def get_acrin_6698_paths( path: Union[os.PathLike, str], max_cases: Optional[int] = None, case_ids: Optional[List[str]] = None, download: bool = False) -> List[str]:
398def get_acrin_6698_paths(
399    path: Union[os.PathLike, str],
400    max_cases: Optional[int] = None,
401    case_ids: Optional[List[str]] = None,
402    download: bool = False,
403) -> List[str]:
404    """Get paths to the ACRIN-6698 data.
405
406    Args:
407        path: Filepath to a folder where the data is downloaded for further processing.
408        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
409        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
410        download: Whether to download the data if it is not present.
411
412    Returns:
413        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
414    """
415    preprocessed_dir = get_acrin_6698_data(path, max_cases, case_ids, download)
416    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
417    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'."
418    return volume_paths

Get paths to the ACRIN-6698 data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • max_cases: The maximum number of cases to use. See get_acrin_6698_data for details.
  • case_ids: Explicit list of case ids to use. See get_acrin_6698_data for details.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').

def get_acrin_6698_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], resize_inputs: bool = False, max_cases: Optional[int] = None, case_ids: Optional[List[str]] = None, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
421def get_acrin_6698_dataset(
422    path: Union[os.PathLike, str],
423    patch_shape: Tuple[int, ...],
424    resize_inputs: bool = False,
425    max_cases: Optional[int] = None,
426    case_ids: Optional[List[str]] = None,
427    download: bool = False,
428    **kwargs
429) -> Dataset:
430    """Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI.
431
432    Args:
433        path: Filepath to a folder where the data is downloaded for further processing.
434        patch_shape: The patch shape to use for training.
435        resize_inputs: Whether to resize inputs to the desired patch shape.
436        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
437        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
438        download: Whether to download the data if it is not present.
439        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
440
441    Returns:
442        The segmentation dataset.
443    """
444    volume_paths = get_acrin_6698_paths(path, max_cases, case_ids, download)
445
446    if resize_inputs:
447        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
448        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
449            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
450        )
451
452    return torch_em.default_segmentation_dataset(
453        raw_paths=volume_paths,
454        raw_key="raw",
455        label_paths=volume_paths,
456        label_key="labels",
457        patch_shape=patch_shape,
458        is_seg_dataset=True,
459        **kwargs
460    )

Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • max_cases: The maximum number of cases to use. See get_acrin_6698_data for details.
  • case_ids: Explicit list of case ids to use. See get_acrin_6698_data for details.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_acrin_6698_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], resize_inputs: bool = False, max_cases: Optional[int] = None, case_ids: Optional[List[str]] = None, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
463def get_acrin_6698_loader(
464    path: Union[os.PathLike, str],
465    batch_size: int,
466    patch_shape: Tuple[int, ...],
467    resize_inputs: bool = False,
468    max_cases: Optional[int] = None,
469    case_ids: Optional[List[str]] = None,
470    download: bool = False,
471    **kwargs
472) -> DataLoader:
473    """Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI.
474
475    Args:
476        path: Filepath to a folder where the data is downloaded for further processing.
477        batch_size: The batch size for training.
478        patch_shape: The patch shape to use for training.
479        resize_inputs: Whether to resize inputs to the desired patch shape.
480        max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details.
481        case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details.
482        download: Whether to download the data if it is not present.
483        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
484
485    Returns:
486        The DataLoader.
487    """
488    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
489    dataset = get_acrin_6698_dataset(path, patch_shape, resize_inputs, max_cases, case_ids, download, **ds_kwargs)
490    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • max_cases: The maximum number of cases to use. See get_acrin_6698_data for details.
  • case_ids: Explicit list of case ids to use. See get_acrin_6698_data for details.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.