torch_em.data.datasets.histopathology.sthelar

The STHELAR dataset contains annotations for nucleus instance segmentation and cell type classification in H&E stained histopathology images of 13 human tissue types (breast, cervix, colon, heart, kidney, liver, lung, lymph node, ovary, pancreas, prostate, skin and tonsil).

The dataset links Xenium spatial transcriptomics with H&E whole slide images. The nuclei masks and the cell types are derived from the spatial transcriptomics data of the slides (cell types via Tangram alignment to single-cell reference atlases, followed by clustering and marker gene based refinement) and are NOT manual annotations, see the publication for the quality control. The slides are cut into patches of 256x256 pixels with an overlap of 64 pixels: the Hugging Face release used here contains 154,814 patches at 20x and 587,555 patches at 40x magnification, extracted from 27 slides (see SLIDES).

Each patch is stored in a separate hdf5 file, which contains the RGB image ('image', channel first) and the labels 'labels/instances' (nucleus instance ids, unique per patch, 0 is background) and 'labels/semantic' (cell type of each nucleus pixel). The cell types are the harmonized 'cells_final_label_group' of the per-slide metadata, with the label ids given by CLASS_NAMES (1-based, 0 is background). Nuclei can be cut at the patch border, and neighbouring patches overlap, so the same cell can appear in several patches.

NOTE: The data is stored in large parquet files, each holding the patches of one or several slides, and this module converts them into the patch files once. The parquet files are kept next to the converted data. All slides need 18 GB (20x) or 54 GB (40x) to download plus the converted patches. Use slides to only download and convert a subset. This requires the pyarrow python package.

The data is located at https://huggingface.co/datasets/FelicieGS/STHELAR_20x and https://huggingface.co/datasets/FelicieGS/STHELAR_40x, released under a CC-BY-4.0 license. The full data, including the per-cell analyses, is available at https://doi.org/10.6019/S-BIAD2146.

This dataset is from the publication https://doi.org/10.1038/s41597-026-06937-6. Please cite it if you use this dataset for your research.

  1"""The STHELAR dataset contains annotations for nucleus instance segmentation and cell type classification
  2in H&E stained histopathology images of 13 human tissue types (breast, cervix, colon, heart, kidney, liver,
  3lung, lymph node, ovary, pancreas, prostate, skin and tonsil).
  4
  5The dataset links Xenium spatial transcriptomics with H&E whole slide images. The nuclei masks and the cell types
  6are derived from the spatial transcriptomics data of the slides (cell types via Tangram alignment to single-cell
  7reference atlases, followed by clustering and marker gene based refinement) and are NOT manual annotations,
  8see the publication for the quality control. The slides are cut into patches of 256x256 pixels with an overlap
  9of 64 pixels: the Hugging Face release used here contains 154,814 patches at 20x and 587,555 patches at 40x
 10magnification, extracted from 27 slides (see `SLIDES`).
 11
 12Each patch is stored in a separate hdf5 file, which contains the RGB image ('image', channel first) and the labels
 13'labels/instances' (nucleus instance ids, unique per patch, 0 is background) and 'labels/semantic' (cell type of
 14each nucleus pixel). The cell types are the harmonized 'cells_final_label_group' of the per-slide metadata, with
 15the label ids given by `CLASS_NAMES` (1-based, 0 is background). Nuclei can be cut at the patch border, and
 16neighbouring patches overlap, so the same cell can appear in several patches.
 17
 18NOTE: The data is stored in large parquet files, each holding the patches of one or several slides, and this
 19module converts them into the patch files once. The parquet files are kept next to the converted data. All slides
 20need 18 GB (20x) or 54 GB (40x) to download plus the converted patches. Use `slides` to only download and convert a
 21subset. This requires the pyarrow python package.
 22
 23The data is located at https://huggingface.co/datasets/FelicieGS/STHELAR_20x and
 24https://huggingface.co/datasets/FelicieGS/STHELAR_40x, released under a CC-BY-4.0 license. The full data,
 25including the per-cell analyses, is available at https://doi.org/10.6019/S-BIAD2146.
 26
 27This dataset is from the publication https://doi.org/10.1038/s41597-026-06937-6.
 28Please cite it if you use this dataset for your research.
 29"""
 30
 31import os
 32import uuid
 33from io import BytesIO
 34from glob import glob
 35from tqdm import tqdm
 36from concurrent import futures
 37from natsort import natsorted
 38from typing import Union, Tuple, List, Literal, Optional, Sequence
 39
 40import numpy as np
 41import imageio.v3 as imageio
 42from scipy.sparse import load_npz
 43
 44from torch.utils.data import Dataset, DataLoader
 45
 46import torch_em
 47
 48from .. import util
 49
 50
 51CLASS_NAMES = [
 52    "Epithelial", "Blood_vessel", "Fibroblast_Myofibroblast", "Myeloid", "B_Plasma", "T_NK", "Melanocyte",
 53    "Specialized", "Other",
 54]
 55CLASS_IDS = {name: i for i, name in enumerate(CLASS_NAMES, start=1)}
 56
 57MAGNIFICATIONS = ["20x", "40x"]
 58COMPLETE_MARKER = ".complete"
 59
 60SLIDES = [
 61    "breast_s0", "breast_s1", "breast_s3", "breast_s6",
 62    "cervix_s0", "colon_s1", "colon_s2", "heart_s0",
 63    "kidney_s0", "kidney_s1", "liver_s0", "liver_s1",
 64    "lung_s1", "lung_s3", "lymph_node_s0", "ovary_s0",
 65    "ovary_s1", "pancreatic_s0", "pancreatic_s1", "pancreatic_s2",
 66    "prostate_s0", "skin_s1", "skin_s2", "skin_s3",
 67    "skin_s4", "tonsil_s0", "tonsil_s1",
 68]
 69
 70REPOS = {"20x": "FelicieGS/STHELAR_20x", "40x": "FelicieGS/STHELAR_40x"}
 71REVISIONS = {
 72    "20x": "a01e22cdb7368edf595edf6150d0ac57deb5e548",
 73    "40x": "e32a8cdd50eff2d38e237f3729e9ac85bbb5203b",
 74}
 75
 76SLIDE_SHARDS = {
 77    "20x": {
 78        "breast_s0": [0, 1],
 79        "breast_s1": [1, 2, 3, 4],
 80        "breast_s3": [4, 5, 6],
 81        "breast_s6": [6, 7],
 82        "cervix_s0": [7, 8],
 83        "colon_s1": [8, 9],
 84        "colon_s2": [9],
 85        "heart_s0": [9],
 86        "kidney_s0": [9],
 87        "kidney_s1": [9, 10],
 88        "liver_s0": [10],
 89        "liver_s1": [10, 11],
 90        "lung_s1": [11],
 91        "lung_s3": [11, 12],
 92        "lymph_node_s0": [12],
 93        "ovary_s0": [12],
 94        "ovary_s1": [12, 13],
 95        "pancreatic_s0": [13, 14],
 96        "pancreatic_s1": [14],
 97        "pancreatic_s2": [14],
 98        "prostate_s0": [14, 15],
 99        "skin_s1": [15],
100        "skin_s2": [15, 16],
101        "skin_s3": [16],
102        "skin_s4": [16],
103        "tonsil_s0": [16, 17],
104        "tonsil_s1": [17],
105    },
106    "40x": {
107        "breast_s0": [0, 1, 2, 3, 4],
108        "breast_s1": [4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15],
109        "breast_s3": [15, 16, 17, 18],
110        "breast_s6": [18, 19, 20, 21, 22],
111        "cervix_s0": [22, 23, 24, 25, 26],
112        "colon_s1": [26, 27],
113        "colon_s2": [27, 28, 29],
114        "heart_s0": [29],
115        "kidney_s0": [29, 30],
116        "kidney_s1": [30],
117        "liver_s0": [30, 31, 32],
118        "liver_s1": [32, 33],
119        "lung_s1": [33, 34],
120        "lung_s3": [34, 35, 36],
121        "lymph_node_s0": [36, 37],
122        "ovary_s0": [37, 38],
123        "ovary_s1": [38, 39, 40, 41],
124        "pancreatic_s0": [41, 42, 43],
125        "pancreatic_s1": [43, 44],
126        "pancreatic_s2": [44, 45],
127        "prostate_s0": [45, 46],
128        "skin_s1": [46, 47],
129        "skin_s2": [47, 48],
130        "skin_s3": [48, 49],
131        "skin_s4": [49],
132        "tonsil_s0": [49, 50, 51, 52],
133        "tonsil_s1": [52, 53],
134    },
135}
136
137SHARD_CHECKSUMS = {
138    "20x": [
139        "2beb26d148a5bd164289708794ebbcdb79b974135732509db8eb5d8a99e61840",
140        "8c4a010b9049a64f93eb3bc819789a713640b815be0b7c7d84b0aff50093a86a",
141        "f1bbcc6b5d3792554e376f5353b9ddc4168167ccb773823dd55bb3967d36cdf9",
142        "e4642280e2a51c3175e2555f2110f88c3c3c448a4c1c0877f20f7b1780c5ba54",
143        "8254d611a92ecd81e8e4f985b352f41767177dc6c941ff4f61e3f9499bce87bb",
144        "64cf474fab05b96ced0cd0e14b429503975ee4e9af472f7d28d07c66c2f43390",
145        "c45fff34cb07e2e28565a8a283b32c032c3c49a6d7b16f819d4970e405e4bcf6",
146        "3201ca4493a73aaef48113127ad33f022c0e3aa878e3ad2935d21ba1f593e86b",
147        "2bdfe7443d7247c89801ea6ad51f2a948b46fd022e1bf33614a8bfc3e9f6c657",
148        "fd0be2332dcda2e95718a245e00943a2289d6fee940a70032d2b3a04bc2110cf",
149        "88cdba926f9487d89e88b09926b2fa1a808bb9f839ef675080980ca262554dfe",
150        "b3de85309f6b2f044b49ada57bbf2e9589293b1da81e436a0c615a945351759e",
151        "cfd75a5137dba283bc2d14cc47f56b1055f30f84a8fc056714f6714fdde1137d",
152        "2dd980ca2cb79ce746fafafdd6ff905c48014734d6bbca0750591a9d22a6a7f5",
153        "3d6b02ceba429b91f122423505f79bcc2c4b7a81ea22dc6b0f63f6b195a2ab63",
154        "051691b8a73b811c929c19cbcfd0611e36c3cfee585d8c837201e6651be9f5c9",
155        "99d9d589bc21788516237414a136c76758a76099713f3eea08d9b6df034bedd8",
156        "d17ba1032ff76fae8d8a9c8ba529facd073519b81423e686790ad6fba773c0eb",
157    ],
158    "40x": [
159        "fded6f9036cdba4c0550e995d1dc2cd46f68ba14276419378d49c523a2ace915",
160        "9090da789118603f02cc451ca4e28a4e0643c447e0a6f61e1f79925022a3a875",
161        "c9ed431b0d63631e299c9fb909812b4a0f51160d4a780cb5f9e5611a3e1265b5",
162        "3e548fa7f80d8ee8455fe399f15eb554954fa987de0348764c0c5f9dac97373c",
163        "2cb12c3eab945479624e93658fc54e6ba441d3eaf279401f440fc2943291f319",
164        "7c198b860fac69fc952fcc42a03ced84e414a4c6618a8322e4c4092ae93f2185",
165        "742ff210dd58adc96e9759037811ddeefd865f042b01e8206e784045dc1868f6",
166        "8263ebaee38fdce69d6a052c9488043b952136e93a8abe7c3852171c382c3cf2",
167        "5b66f363aa6dd86d98eaed356ca16b53df0dcb6cb4ae6d9dc2c72f15ba38865c",
168        "7d39f3cde0a3000b4ab6b0bfb9d845ba14cdb80574fd1f17cb64e0cb8e81a5e9",
169        "90ede649a250ea7877d97b400bbe64f01506b96de5b6e420eb9a5ab08cebd279",
170        "5e79e72a0fa406e0f070c65b47b07a9b390c9c712a055185e0ab34059fa2ed84",
171        "df044c5bfc34be9e3d930d286fd7665e19810ed3d8a42dde7233b0deb8e5b8b3",
172        "57baa4ee89d96e0b4712e874a680a7cd3ce4b92b38385576d9593aadcfa7bf7f",
173        "db5d9801b45ef358d1c8b90865ebcef8660f4628ab5195a9d0651a8373dffada",
174        "1cdd7d08ec094f4f950825af322518558928523a1e58c44e59b6e9d44f8ec6e1",
175        "463c18d8cacd82e05738a00b6bd09bbfc1f53eab93c04abf78df0935afb6c4df",
176        "b2a265d3a65d1a3c1461763fde7a088b4f6d3ed94f5ec6bb04e3debcb84201aa",
177        "b62961040f0ba1e5ac56dc3a328a1b478001bee668e03b03fde4a75cc2cf7266",
178        "354f00172234bd2e788b7874c31bab201eff84fafaf19ad425e2d8705b79b8d6",
179        "90a20a79276eb571d8840ee3391923eca32d075c940926a674a51f8fa86a5cf8",
180        "38601ae47fa7ac50a6327e466395acb82f20adb1a18314d578f11458f2d55574",
181        "96fbf1edb41689439d7e327dab80a24a6b720f143054b35de1193d21c49cad14",
182        "889e24ea440f9bbc0e086f91086e4f30a6a409c0aa4ab4004ec8887a7897e3ea",
183        "38dc9cf1a2bd50c6b21e32f4fc2dfc630d0691acee6e17159ff2f741bcd36379",
184        "19f8bd33957c0fa13e54f858a8ea422095ee4ad30708c9f872c66f81ec7f135a",
185        "a478b13ae480f546f3f74df8af6e5e7be78a83d82574814d3e4994205e317ebb",
186        "b7c6a826ebfa182c421c652fe850c2efc0d9e23090ea6378be09925297197862",
187        "042eef94be71a40268f0fd12a88bded6d1d99ec08aace18233728823a61523de",
188        "c1308467499672cd49110a7ddea1bb0ce0d009faba744da67a6da0e305d93bb5",
189        "7ba700b6d595e7e6f72103392ebc16cd1cdf94c79174ae79765cf96768e8bddc",
190        "ad928308055f7e86813ab613a05aa0c6bd65f78cfe7bab8892c38f90264fbba9",
191        "780810e23adac20be6e5c8fd503e4261c5cf78dfd53ce3e36056294288a6dd44",
192        "89f577a75a07026fb3293fcac5a9d451958232a0960866908c5282f9eaae5817",
193        "8df8f7310cbba63e047e8ae9d30987e09f4697ae927cc7c7b7e08d22bcf093f6",
194        "aaf8357b174e2a53532bcb213d3ba21e79dfa4c82cf39ebf338bb6351f1f1de6",
195        "6301ba19059423d481e11617237c058f9b8382f4993401baa14ffc9732d6e3c6",
196        "289e3dbbaaf77e06a57e4aa711f61f5eb4d7fb6b3f97247967874400db64ea57",
197        "5d67d1601a292dc9f371f2ea3349036fde6bbcb18eaa035349c3600ec0dbeb31",
198        "ec7b89e7acd33c9dfbc6b9e12737bcb6da887fd2ca8634cb330b32455c4c5833",
199        "cacb6d40d9dd4e9fe0441c18c0d2ed687da738fc7c12c1767feaa96ad6a52f3d",
200        "571a72273be1fa6f3c14e99087ba130f2f83b098f9674291e1f705f0e4e2752d",
201        "1470a4745b0fe58253d7c484e4d6304e9fc3c008158f66fe403f6fb00e698ee4",
202        "f33e597a2aba5eff981144c5eaa89d51844de76a1a9dcf58b22da7e9a6a301a4",
203        "42e617f6f15ce018d2d11a5ba20527c97414aaea14debf28600240542cf55f9c",
204        "000ecc178dab26dcca21a282d410b51b2b3bec8edb800edc60baa40fd34124b7",
205        "2240e03cfd020c8a9e53524e335f17d1edb4b28319584726efd77429587e9547",
206        "487f7a86152e7acfd623a27fade0fc0c88ad89367f764971bba1f619bff5e140",
207        "af125205e4ff222db33611a2de2ef398e485f11ad2e492632ecb6524de0a0b07",
208        "67ea31dc49c0022d803e143c41e79a61feca2bfb98d0e001c45523b33741d4eb",
209        "d0883ae9c68b7eafad1ae9b8b75a3befcc1aaa86bb1120d022b90854e9d73ed8",
210        "10285d36a9024b2a036b3f019f294dcc6a38d41b348a8176702a7ce87c587f02",
211        "d28f98f0bf468f2fc14f95f1c721175513269ee9a18a13d4314c5e5a848d2e23",
212        "94f2cea4dfd344933c942d14243744b882cc2c10d42628f95ff51c7253c28df7",
213    ],
214}
215
216METADATA_CHECKSUMS = {
217    "20x": {
218        "breast_s0": "1cdb4bba93afb1c3d3eeca8f2caacb6a5d10b22f07b51acce274f8f015588bf5",
219        "breast_s1": "abf1ebba18841e7a86f3114516a07574221f7d6978989c232dedcc7a258169e6",
220        "breast_s3": "435e1a77a0acb2842d08276ad7e3cc67916a6e3949a56aca2260b831b5e0e26e",
221        "breast_s6": "02b60f4bcb26646bc9273acd0dfc6ee98f6a6811266c7b862a625c2f5d87c15d",
222        "cervix_s0": "3a198703ca5e228c81ea03166082316f857cc42416da5b303090206ca24abb4e",
223        "colon_s1": "0f11618dae9209537855ce9568139181cc7d6223ea98810907ad299e77a0067e",
224        "colon_s2": "823b0be5277fabf4930eae5d4402f4fabafd8876dc674c19e641b203a20a96b9",
225        "heart_s0": "4f1a38ab08f83603faac675a3fba49fcc3528fda91a5c288b4efd5d2f34a878a",
226        "kidney_s0": "a4db038e375a8f47532f689fd457516e761acf9c23cb90e128ac5d134b1b4f20",
227        "kidney_s1": "a9bf51936b349456b836c80de6ba131852e91e3e040f20eda3802d837a27533c",
228        "liver_s0": "17d9a987a88460824dabde15a55e5e864cf606aed57e2d3f2f8b4505b9c2d2eb",
229        "liver_s1": "1e50bcda0c31cdd25a84bed02d6c548881a3f3e8a20b57d3daa4982dad505bc1",
230        "lung_s1": "83b059126e8a1bac93eee91cd36ffda90e02aac406f241dc17b09963bf488463",
231        "lung_s3": "639ded2b7c195a44ad7708fef31e4314505fdced5a8097e0e95a3ab5624cd361",
232        "lymph_node_s0": "040c24a0bcd156ecbdfb13d476bae127bf4339cd189dae460ee42f7c4bb0242a",
233        "ovary_s0": "b77a2d78eaff3a69bc4796ab9310c1f708b0368ceaaeb7a3196490e4bbc5160f",
234        "ovary_s1": "2e66104150da680f19789a1e6a97b06f0e5710c83a2803a6754278838c8c59ad",
235        "pancreatic_s0": "5b242492f5ddd965cfa8e6fa991ceea51b9d549fccfb56e91c121e46baba7a6f",
236        "pancreatic_s1": "a1f37d4b57028f814be2f55ff31f74c87f1e82f482558071086ae79078d8dd0e",
237        "pancreatic_s2": "f006c40cbf16a5a2ccf0782fe872f7150d42c5247fec5e93452ae74e132cac56",
238        "prostate_s0": "e719328ca0239dc0ec5f252f2dea88e5c0e94595b7911f4759c1df935bfbd114",
239        "skin_s1": "60354b20b2824abea0dfec2c70e4df2555329d6c2672091ce47e07a88f038af5",
240        "skin_s2": "ef4cd89a2df0ef04f777a184667661f1266a0e4cb6a2fc633e57f84b1783f95c",
241        "skin_s3": "150412b1a9ff4489756567f970a9663b5c894fd50093241316f5e72ef7008e49",
242        "skin_s4": "df44e89d5efbdbfb5d3702e92921ebc38060ee5df02bbb8a87604467780013ea",
243        "tonsil_s0": "8edd68ce0858c303386ed017264b89d7ffa90ba61dec9cc937ed849668394493",
244        "tonsil_s1": "3dc4adb307b85ac7e53bcf65a8d77ce02b62f058d58fcab8543618e804f815e9",
245    },
246    "40x": {
247        "breast_s0": "f8ddf4696e0a0dd927a8b0c1e24242d3f4575190c7a26fe80f78d0932bb23227",
248        "breast_s1": "5fea33e2716a92054f02f3d8d75ca96a37f7e75acf91821451f6bd771d2fbc00",
249        "breast_s3": "457a3fdabaad46f8e007958419a5828a6d50e438f236b1bb0064b8d604712730",
250        "breast_s6": "64036da240ccf08273156939dbe1a3545951d07d1f62d97db7d3a3f7c8758689",
251        "cervix_s0": "dc980ba27a43a6fe826e02cf529cf17e35f98a22781a530cd5c2b298ebb89d1c",
252        "colon_s1": "27a31ef22fe716285f63dab79bf45aa9cdf6e8e860b8ca725910a239b800050c",
253        "colon_s2": "ad1d3e2cfa20c39da3c4db03c4760732a0e4d851b3ef7f8b45f30cb6ba74756f",
254        "heart_s0": "cfa6690e8be28714583edb866cb3fadf575733cdd084b6d3e8cc68e49b3917d8",
255        "kidney_s0": "70ea019dc4c89647fe4ea34f65061b1386cac46cf69f78e036b1d3f14b6c1766",
256        "kidney_s1": "88fd5484af6f5d1734a7f0f7a8fe6e50151a39c73c48a11892baf79ffd1f724e",
257        "liver_s0": "b6f74e1ecfaa59bbb75921ba6855877f5692487e6890e7a396cd0684d7cb467d",
258        "liver_s1": "a6058b2e2ed771e7cf82dfc1f22c3590fe99d003abdd22d18cbe9266f44bc03d",
259        "lung_s1": "041fbcae0f18ae626c2b5daa37231020bad1edf589ea5149418ca9b727eae2c8",
260        "lung_s3": "4c2ad9337adb0683604f0c8683288cb8eb902f095dc09eeb8764c58ab6c75386",
261        "lymph_node_s0": "e1b4659526ee64a22eada40cec267a7a1b75fc3079d84f6123bfe547bf2716b2",
262        "ovary_s0": "9e561f24a1cc9b6a3dfd9e2c2320c5d11a9ebae9d54bc5ddd6d9db3c446b5375",
263        "ovary_s1": "79f100d35cedee540d5f5319549a381faf8ebde736601e329a99f815b3978980",
264        "pancreatic_s0": "f154d79606c8d1acff94c032bfc975e8518913f488da2bc9224d8ae32162c05b",
265        "pancreatic_s1": "e46c59592eea273450012d599dc09e87a970dab35eae1e9f00f431ecdd1201ed",
266        "pancreatic_s2": "efccb9598ad581bcc0da9fd7c4b2b585b7c405774f0359e2380dd2ae51f56a54",
267        "prostate_s0": "5397a52b30862f60780a8d105cacfa5a5ef7db66896edc88def2a6aa7b55cb3e",
268        "skin_s1": "4b4faf5248e8555fd68395f16d7a6ca653c02f7b45171da224f77026e5246f8c",
269        "skin_s2": "18b149fae149ac1b35c09701d8815b5818c90b5670643269a465d6aa34597b62",
270        "skin_s3": "2cd2e0c9d49c56192cb2eec65d2f4bec3feb0fef3a986472a2c50113e61e5bd5",
271        "skin_s4": "e5879ec0348b02eb1c6cb8ca04b7c33f88d50328e95008309453ef24e21e1702",
272        "tonsil_s0": "0f95581fdc179715ac190df5e6a02df03e48baf156c4c700fc011101ba9315ff",
273        "tonsil_s1": "5f40a96f6c11cdfba8f80f4eb0bdf14c0b34000166c5d116a3714bf25e478222",
274    },
275}
276
277
278def _validate(magnification, slides):
279    if magnification not in MAGNIFICATIONS:
280        raise ValueError(f"'{magnification}' is not a valid magnification. Choose one of {MAGNIFICATIONS}.")
281    if slides is None:
282        return list(SLIDES)
283    unknown = [s for s in slides if s not in SLIDES]
284    if unknown:
285        raise ValueError(f"{unknown} are not valid slides. Choose from {SLIDES}.")
286    return [s for s in SLIDES if s in slides]
287
288
289def _shard_name(magnification, shard_id):
290    return f"train-{shard_id:05d}-of-{len(SHARD_CHECKSUMS[magnification]):05d}.parquet"
291
292
293def _resolve_url(magnification, relative_path):
294    return f"https://huggingface.co/datasets/{REPOS[magnification]}/resolve/{REVISIONS[magnification]}/{relative_path}"
295
296
297def _load_lookup(metadata_path):
298    import pandas as pd
299
300    df = pd.read_parquet(metadata_path, columns=["cell_id_int", "cells_final_label_group"])
301    codes = df["cells_final_label_group"].map(CLASS_IDS)
302    assert not codes.isna().any(), f"Unknown cell types in {metadata_path}."
303    ids = df["cell_id_int"].to_numpy().astype("int64")
304    order = np.argsort(ids)
305    return ids[order], codes.to_numpy().astype("uint8")[order]
306
307
308def _write_patch(out_path, image_bytes, cell_id_bytes, lookup):
309    if os.path.exists(out_path):
310        return
311
312    import h5py
313
314    image = imageio.imread(image_bytes, extension=".png")
315    assert image.ndim == 3 and image.shape[-1] == 3
316    cell_ids = load_npz(BytesIO(cell_id_bytes)).toarray().astype("int64")
317    assert cell_ids.shape == image.shape[:2]
318
319    unique_ids, inverse = np.unique(cell_ids, return_inverse=True)
320    instances = inverse.reshape(cell_ids.shape)
321    if unique_ids[0] == 0:
322        nucleus_ids = unique_ids[1:]
323    else:
324        nucleus_ids = unique_ids
325        instances = instances + 1
326
327    meta_ids, meta_codes = lookup
328    positions = np.minimum(np.searchsorted(meta_ids, nucleus_ids), len(meta_ids) - 1)
329    assert (meta_ids[positions] == nucleus_ids).all(), "Found cell ids without metadata."
330    semantic = np.concatenate([[0], meta_codes[positions]]).astype("uint8")[instances]
331
332    tmp_path = f"{out_path}.{uuid.uuid4().hex}.incomplete"
333    with h5py.File(tmp_path, "w") as f:
334        f.create_dataset("image", data=image.transpose((2, 0, 1)), compression="gzip")
335        f.create_dataset("labels/instances", data=instances.astype("int32"), compression="gzip")
336        f.create_dataset("labels/semantic", data=semantic, compression="gzip")
337    os.replace(tmp_path, out_path)
338
339
340def _convert_shard(shard_path, slides, lookups, preprocessed_dir, n_workers):
341    import pyarrow.parquet as pq
342
343    parquet_file = pq.ParquetFile(shard_path)
344    with futures.ThreadPoolExecutor(n_workers) as pool:
345        for row_group in tqdm(range(parquet_file.num_row_groups), desc=f"Convert {os.path.basename(shard_path)}"):
346            table = parquet_file.read_row_group(row_group, columns=["file_name", "slide_id", "image", "cell_id_map"])
347            slide_ids = table.column("slide_id").to_pylist()
348            rows = [i for i, slide_id in enumerate(slide_ids) if slide_id in slides]
349            if not rows:
350                continue
351
352            file_names = table.column("file_name").to_pylist()
353            images = table.column("image").to_pylist()
354            cell_id_maps = table.column("cell_id_map").to_pylist()
355            tasks = []
356            for i in rows:
357                out_path = os.path.join(preprocessed_dir, slide_ids[i], f"{os.path.splitext(file_names[i])[0]}.h5")
358                tasks.append(
359                    pool.submit(_write_patch, out_path, images[i]["bytes"], cell_id_maps[i], lookups[slide_ids[i]])
360                )
361            for task in tasks:
362                task.result()
363
364
365def get_sthelar_data(
366    path: Union[os.PathLike, str],
367    magnification: Literal["20x", "40x"] = "20x",
368    slides: Optional[Sequence[str]] = None,
369    download: bool = False,
370) -> str:
371    """Download the STHELAR dataset and convert its patches to hdf5 files.
372
373    NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use `slides` for a subset. The parquet files
374    of the slides are kept, since a parquet file can contain the patches of several slides.
375
376    Args:
377        path: Filepath to a folder where the data is downloaded for further processing.
378        magnification: The choice of magnification. Either '20x' or '40x'.
379        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
380        download: Whether to download the data if it is not present.
381
382    Returns:
383        Filepath where the converted patches are stored, in one sub-folder per slide.
384    """
385    slides = _validate(magnification, slides)
386    data_dir = os.path.join(path, magnification)
387    preprocessed_dir = os.path.join(data_dir, "preprocessed")
388
389    pending = [s for s in slides if not os.path.exists(os.path.join(preprocessed_dir, s, COMPLETE_MARKER))]
390    if not pending:
391        return preprocessed_dir
392
393    os.makedirs(os.path.join(data_dir, "data"), exist_ok=True)
394    os.makedirs(os.path.join(data_dir, "cell_metadata"), exist_ok=True)
395
396    lookups = {}
397    for slide in pending:
398        relative_path = f"cell_metadata/{slide}_cell_metadata.parquet"
399        metadata_path = os.path.join(data_dir, relative_path)
400        util.download_source(
401            path=metadata_path, url=_resolve_url(magnification, relative_path), download=download,
402            checksum=METADATA_CHECKSUMS[magnification][slide],
403        )
404        lookups[slide] = _load_lookup(metadata_path)
405        os.makedirs(os.path.join(preprocessed_dir, slide), exist_ok=True)
406
407    n_cpus = len(os.sched_getaffinity(0)) if hasattr(os, "sched_getaffinity") else (os.cpu_count() or 1)
408    n_workers = min(8, n_cpus)
409
410    shard_ids = sorted({i for slide in pending for i in SLIDE_SHARDS[magnification][slide]})
411    for shard_id in shard_ids:
412        relative_path = f"data/{_shard_name(magnification, shard_id)}"
413        shard_path = os.path.join(data_dir, relative_path)
414        util.download_source(
415            path=shard_path, url=_resolve_url(magnification, relative_path), download=download,
416            checksum=SHARD_CHECKSUMS[magnification][shard_id],
417        )
418        _convert_shard(shard_path, set(pending), lookups, preprocessed_dir, n_workers)
419
420    for slide in pending:
421        with open(os.path.join(preprocessed_dir, slide, COMPLETE_MARKER), "w"):
422            pass
423
424    return preprocessed_dir
425
426
427def get_sthelar_paths(
428    path: Union[os.PathLike, str],
429    magnification: Literal["20x", "40x"] = "20x",
430    slides: Optional[Sequence[str]] = None,
431    download: bool = False,
432) -> List[str]:
433    """Get paths to the STHELAR data.
434
435    Args:
436        path: Filepath to a folder where the data is downloaded for further processing.
437        magnification: The choice of magnification. Either '20x' or '40x'.
438        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
439        download: Whether to download the data if it is not present.
440
441    Returns:
442        List of filepaths for the hdf5 files, which contain the image ('image') and the labels
443        ('labels/instances' and 'labels/semantic') of a patch.
444    """
445    slides = _validate(magnification, slides)
446    preprocessed_dir = get_sthelar_data(path, magnification, slides, download)
447
448    data_paths = []
449    for slide in slides:
450        data_paths.extend(natsorted(glob(os.path.join(preprocessed_dir, slide, "*.h5"))))
451
452    assert len(data_paths) > 0
453    return data_paths
454
455
456def get_sthelar_dataset(
457    path: Union[os.PathLike, str],
458    patch_shape: Tuple[int, int],
459    magnification: Literal["20x", "40x"] = "20x",
460    slides: Optional[Sequence[str]] = None,
461    label_type: Literal["instances", "semantic"] = "instances",
462    resize_inputs: bool = False,
463    download: bool = False,
464    **kwargs
465) -> Dataset:
466    """Get the STHELAR dataset for nucleus segmentation and cell type classification.
467
468    Args:
469        path: Filepath to a folder where the data is downloaded for further processing.
470        patch_shape: The patch shape to use for training.
471        magnification: The choice of magnification. Either '20x' or '40x'.
472        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
473        label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
474            cell types of the nuclei, with the ids given by `CLASS_NAMES`.
475        resize_inputs: Whether to resize the inputs to the patch shape.
476        download: Whether to download the data if it is not present.
477        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
478
479    Returns:
480        The segmentation dataset.
481    """
482    if label_type not in ("instances", "semantic"):
483        raise ValueError(f"'{label_type}' is not a valid label type. Choose either 'instances' or 'semantic'.")
484
485    data_paths = get_sthelar_paths(path, magnification, slides, download)
486
487    if resize_inputs:
488        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
489        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
490            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
491        )
492
493    return torch_em.default_segmentation_dataset(
494        raw_paths=data_paths,
495        raw_key="image",
496        label_paths=data_paths,
497        label_key=f"labels/{label_type}",
498        patch_shape=patch_shape,
499        ndim=2,
500        with_channels=True,
501        **kwargs
502    )
503
504
505def get_sthelar_loader(
506    path: Union[os.PathLike, str],
507    batch_size: int,
508    patch_shape: Tuple[int, int],
509    magnification: Literal["20x", "40x"] = "20x",
510    slides: Optional[Sequence[str]] = None,
511    label_type: Literal["instances", "semantic"] = "instances",
512    resize_inputs: bool = False,
513    download: bool = False,
514    **kwargs
515) -> DataLoader:
516    """Get the STHELAR dataloader for nucleus segmentation and cell type classification.
517
518    Args:
519        path: Filepath to a folder where the data is downloaded for further processing.
520        batch_size: The batch size for training.
521        patch_shape: The patch shape to use for training.
522        magnification: The choice of magnification. Either '20x' or '40x'.
523        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
524        label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
525            cell types of the nuclei, with the ids given by `CLASS_NAMES`.
526        resize_inputs: Whether to resize the inputs to the patch shape.
527        download: Whether to download the data if it is not present.
528        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
529
530    Returns:
531        The DataLoader.
532    """
533    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
534    dataset = get_sthelar_dataset(
535        path, patch_shape, magnification, slides, label_type, resize_inputs, download, **ds_kwargs
536    )
537    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
CLASS_NAMES = ['Epithelial', 'Blood_vessel', 'Fibroblast_Myofibroblast', 'Myeloid', 'B_Plasma', 'T_NK', 'Melanocyte', 'Specialized', 'Other']
CLASS_IDS = {'Epithelial': 1, 'Blood_vessel': 2, 'Fibroblast_Myofibroblast': 3, 'Myeloid': 4, 'B_Plasma': 5, 'T_NK': 6, 'Melanocyte': 7, 'Specialized': 8, 'Other': 9}
MAGNIFICATIONS = ['20x', '40x']
COMPLETE_MARKER = '.complete'
SLIDES = ['breast_s0', 'breast_s1', 'breast_s3', 'breast_s6', 'cervix_s0', 'colon_s1', 'colon_s2', 'heart_s0', 'kidney_s0', 'kidney_s1', 'liver_s0', 'liver_s1', 'lung_s1', 'lung_s3', 'lymph_node_s0', 'ovary_s0', 'ovary_s1', 'pancreatic_s0', 'pancreatic_s1', 'pancreatic_s2', 'prostate_s0', 'skin_s1', 'skin_s2', 'skin_s3', 'skin_s4', 'tonsil_s0', 'tonsil_s1']
REPOS = {'20x': 'FelicieGS/STHELAR_20x', '40x': 'FelicieGS/STHELAR_40x'}
REVISIONS = {'20x': 'a01e22cdb7368edf595edf6150d0ac57deb5e548', '40x': 'e32a8cdd50eff2d38e237f3729e9ac85bbb5203b'}
SLIDE_SHARDS = {'20x': {'breast_s0': [0, 1], 'breast_s1': [1, 2, 3, 4], 'breast_s3': [4, 5, 6], 'breast_s6': [6, 7], 'cervix_s0': [7, 8], 'colon_s1': [8, 9], 'colon_s2': [9], 'heart_s0': [9], 'kidney_s0': [9], 'kidney_s1': [9, 10], 'liver_s0': [10], 'liver_s1': [10, 11], 'lung_s1': [11], 'lung_s3': [11, 12], 'lymph_node_s0': [12], 'ovary_s0': [12], 'ovary_s1': [12, 13], 'pancreatic_s0': [13, 14], 'pancreatic_s1': [14], 'pancreatic_s2': [14], 'prostate_s0': [14, 15], 'skin_s1': [15], 'skin_s2': [15, 16], 'skin_s3': [16], 'skin_s4': [16], 'tonsil_s0': [16, 17], 'tonsil_s1': [17]}, '40x': {'breast_s0': [0, 1, 2, 3, 4], 'breast_s1': [4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15], 'breast_s3': [15, 16, 17, 18], 'breast_s6': [18, 19, 20, 21, 22], 'cervix_s0': [22, 23, 24, 25, 26], 'colon_s1': [26, 27], 'colon_s2': [27, 28, 29], 'heart_s0': [29], 'kidney_s0': [29, 30], 'kidney_s1': [30], 'liver_s0': [30, 31, 32], 'liver_s1': [32, 33], 'lung_s1': [33, 34], 'lung_s3': [34, 35, 36], 'lymph_node_s0': [36, 37], 'ovary_s0': [37, 38], 'ovary_s1': [38, 39, 40, 41], 'pancreatic_s0': [41, 42, 43], 'pancreatic_s1': [43, 44], 'pancreatic_s2': [44, 45], 'prostate_s0': [45, 46], 'skin_s1': [46, 47], 'skin_s2': [47, 48], 'skin_s3': [48, 49], 'skin_s4': [49], 'tonsil_s0': [49, 50, 51, 52], 'tonsil_s1': [52, 53]}}
SHARD_CHECKSUMS = {'20x': ['2beb26d148a5bd164289708794ebbcdb79b974135732509db8eb5d8a99e61840', '8c4a010b9049a64f93eb3bc819789a713640b815be0b7c7d84b0aff50093a86a', 'f1bbcc6b5d3792554e376f5353b9ddc4168167ccb773823dd55bb3967d36cdf9', 'e4642280e2a51c3175e2555f2110f88c3c3c448a4c1c0877f20f7b1780c5ba54', '8254d611a92ecd81e8e4f985b352f41767177dc6c941ff4f61e3f9499bce87bb', '64cf474fab05b96ced0cd0e14b429503975ee4e9af472f7d28d07c66c2f43390', 'c45fff34cb07e2e28565a8a283b32c032c3c49a6d7b16f819d4970e405e4bcf6', '3201ca4493a73aaef48113127ad33f022c0e3aa878e3ad2935d21ba1f593e86b', '2bdfe7443d7247c89801ea6ad51f2a948b46fd022e1bf33614a8bfc3e9f6c657', 'fd0be2332dcda2e95718a245e00943a2289d6fee940a70032d2b3a04bc2110cf', '88cdba926f9487d89e88b09926b2fa1a808bb9f839ef675080980ca262554dfe', 'b3de85309f6b2f044b49ada57bbf2e9589293b1da81e436a0c615a945351759e', 'cfd75a5137dba283bc2d14cc47f56b1055f30f84a8fc056714f6714fdde1137d', '2dd980ca2cb79ce746fafafdd6ff905c48014734d6bbca0750591a9d22a6a7f5', '3d6b02ceba429b91f122423505f79bcc2c4b7a81ea22dc6b0f63f6b195a2ab63', '051691b8a73b811c929c19cbcfd0611e36c3cfee585d8c837201e6651be9f5c9', '99d9d589bc21788516237414a136c76758a76099713f3eea08d9b6df034bedd8', 'd17ba1032ff76fae8d8a9c8ba529facd073519b81423e686790ad6fba773c0eb'], '40x': ['fded6f9036cdba4c0550e995d1dc2cd46f68ba14276419378d49c523a2ace915', '9090da789118603f02cc451ca4e28a4e0643c447e0a6f61e1f79925022a3a875', 'c9ed431b0d63631e299c9fb909812b4a0f51160d4a780cb5f9e5611a3e1265b5', '3e548fa7f80d8ee8455fe399f15eb554954fa987de0348764c0c5f9dac97373c', '2cb12c3eab945479624e93658fc54e6ba441d3eaf279401f440fc2943291f319', '7c198b860fac69fc952fcc42a03ced84e414a4c6618a8322e4c4092ae93f2185', '742ff210dd58adc96e9759037811ddeefd865f042b01e8206e784045dc1868f6', '8263ebaee38fdce69d6a052c9488043b952136e93a8abe7c3852171c382c3cf2', '5b66f363aa6dd86d98eaed356ca16b53df0dcb6cb4ae6d9dc2c72f15ba38865c', '7d39f3cde0a3000b4ab6b0bfb9d845ba14cdb80574fd1f17cb64e0cb8e81a5e9', '90ede649a250ea7877d97b400bbe64f01506b96de5b6e420eb9a5ab08cebd279', '5e79e72a0fa406e0f070c65b47b07a9b390c9c712a055185e0ab34059fa2ed84', 'df044c5bfc34be9e3d930d286fd7665e19810ed3d8a42dde7233b0deb8e5b8b3', '57baa4ee89d96e0b4712e874a680a7cd3ce4b92b38385576d9593aadcfa7bf7f', 'db5d9801b45ef358d1c8b90865ebcef8660f4628ab5195a9d0651a8373dffada', '1cdd7d08ec094f4f950825af322518558928523a1e58c44e59b6e9d44f8ec6e1', '463c18d8cacd82e05738a00b6bd09bbfc1f53eab93c04abf78df0935afb6c4df', 'b2a265d3a65d1a3c1461763fde7a088b4f6d3ed94f5ec6bb04e3debcb84201aa', 'b62961040f0ba1e5ac56dc3a328a1b478001bee668e03b03fde4a75cc2cf7266', '354f00172234bd2e788b7874c31bab201eff84fafaf19ad425e2d8705b79b8d6', '90a20a79276eb571d8840ee3391923eca32d075c940926a674a51f8fa86a5cf8', '38601ae47fa7ac50a6327e466395acb82f20adb1a18314d578f11458f2d55574', '96fbf1edb41689439d7e327dab80a24a6b720f143054b35de1193d21c49cad14', '889e24ea440f9bbc0e086f91086e4f30a6a409c0aa4ab4004ec8887a7897e3ea', '38dc9cf1a2bd50c6b21e32f4fc2dfc630d0691acee6e17159ff2f741bcd36379', '19f8bd33957c0fa13e54f858a8ea422095ee4ad30708c9f872c66f81ec7f135a', 'a478b13ae480f546f3f74df8af6e5e7be78a83d82574814d3e4994205e317ebb', 'b7c6a826ebfa182c421c652fe850c2efc0d9e23090ea6378be09925297197862', '042eef94be71a40268f0fd12a88bded6d1d99ec08aace18233728823a61523de', 'c1308467499672cd49110a7ddea1bb0ce0d009faba744da67a6da0e305d93bb5', '7ba700b6d595e7e6f72103392ebc16cd1cdf94c79174ae79765cf96768e8bddc', 'ad928308055f7e86813ab613a05aa0c6bd65f78cfe7bab8892c38f90264fbba9', '780810e23adac20be6e5c8fd503e4261c5cf78dfd53ce3e36056294288a6dd44', '89f577a75a07026fb3293fcac5a9d451958232a0960866908c5282f9eaae5817', '8df8f7310cbba63e047e8ae9d30987e09f4697ae927cc7c7b7e08d22bcf093f6', 'aaf8357b174e2a53532bcb213d3ba21e79dfa4c82cf39ebf338bb6351f1f1de6', '6301ba19059423d481e11617237c058f9b8382f4993401baa14ffc9732d6e3c6', '289e3dbbaaf77e06a57e4aa711f61f5eb4d7fb6b3f97247967874400db64ea57', '5d67d1601a292dc9f371f2ea3349036fde6bbcb18eaa035349c3600ec0dbeb31', 'ec7b89e7acd33c9dfbc6b9e12737bcb6da887fd2ca8634cb330b32455c4c5833', 'cacb6d40d9dd4e9fe0441c18c0d2ed687da738fc7c12c1767feaa96ad6a52f3d', '571a72273be1fa6f3c14e99087ba130f2f83b098f9674291e1f705f0e4e2752d', '1470a4745b0fe58253d7c484e4d6304e9fc3c008158f66fe403f6fb00e698ee4', 'f33e597a2aba5eff981144c5eaa89d51844de76a1a9dcf58b22da7e9a6a301a4', '42e617f6f15ce018d2d11a5ba20527c97414aaea14debf28600240542cf55f9c', '000ecc178dab26dcca21a282d410b51b2b3bec8edb800edc60baa40fd34124b7', '2240e03cfd020c8a9e53524e335f17d1edb4b28319584726efd77429587e9547', '487f7a86152e7acfd623a27fade0fc0c88ad89367f764971bba1f619bff5e140', 'af125205e4ff222db33611a2de2ef398e485f11ad2e492632ecb6524de0a0b07', '67ea31dc49c0022d803e143c41e79a61feca2bfb98d0e001c45523b33741d4eb', 'd0883ae9c68b7eafad1ae9b8b75a3befcc1aaa86bb1120d022b90854e9d73ed8', '10285d36a9024b2a036b3f019f294dcc6a38d41b348a8176702a7ce87c587f02', 'd28f98f0bf468f2fc14f95f1c721175513269ee9a18a13d4314c5e5a848d2e23', '94f2cea4dfd344933c942d14243744b882cc2c10d42628f95ff51c7253c28df7']}
METADATA_CHECKSUMS = {'20x': {'breast_s0': '1cdb4bba93afb1c3d3eeca8f2caacb6a5d10b22f07b51acce274f8f015588bf5', 'breast_s1': 'abf1ebba18841e7a86f3114516a07574221f7d6978989c232dedcc7a258169e6', 'breast_s3': '435e1a77a0acb2842d08276ad7e3cc67916a6e3949a56aca2260b831b5e0e26e', 'breast_s6': '02b60f4bcb26646bc9273acd0dfc6ee98f6a6811266c7b862a625c2f5d87c15d', 'cervix_s0': '3a198703ca5e228c81ea03166082316f857cc42416da5b303090206ca24abb4e', 'colon_s1': '0f11618dae9209537855ce9568139181cc7d6223ea98810907ad299e77a0067e', 'colon_s2': '823b0be5277fabf4930eae5d4402f4fabafd8876dc674c19e641b203a20a96b9', 'heart_s0': '4f1a38ab08f83603faac675a3fba49fcc3528fda91a5c288b4efd5d2f34a878a', 'kidney_s0': 'a4db038e375a8f47532f689fd457516e761acf9c23cb90e128ac5d134b1b4f20', 'kidney_s1': 'a9bf51936b349456b836c80de6ba131852e91e3e040f20eda3802d837a27533c', 'liver_s0': '17d9a987a88460824dabde15a55e5e864cf606aed57e2d3f2f8b4505b9c2d2eb', 'liver_s1': '1e50bcda0c31cdd25a84bed02d6c548881a3f3e8a20b57d3daa4982dad505bc1', 'lung_s1': '83b059126e8a1bac93eee91cd36ffda90e02aac406f241dc17b09963bf488463', 'lung_s3': '639ded2b7c195a44ad7708fef31e4314505fdced5a8097e0e95a3ab5624cd361', 'lymph_node_s0': '040c24a0bcd156ecbdfb13d476bae127bf4339cd189dae460ee42f7c4bb0242a', 'ovary_s0': 'b77a2d78eaff3a69bc4796ab9310c1f708b0368ceaaeb7a3196490e4bbc5160f', 'ovary_s1': '2e66104150da680f19789a1e6a97b06f0e5710c83a2803a6754278838c8c59ad', 'pancreatic_s0': '5b242492f5ddd965cfa8e6fa991ceea51b9d549fccfb56e91c121e46baba7a6f', 'pancreatic_s1': 'a1f37d4b57028f814be2f55ff31f74c87f1e82f482558071086ae79078d8dd0e', 'pancreatic_s2': 'f006c40cbf16a5a2ccf0782fe872f7150d42c5247fec5e93452ae74e132cac56', 'prostate_s0': 'e719328ca0239dc0ec5f252f2dea88e5c0e94595b7911f4759c1df935bfbd114', 'skin_s1': '60354b20b2824abea0dfec2c70e4df2555329d6c2672091ce47e07a88f038af5', 'skin_s2': 'ef4cd89a2df0ef04f777a184667661f1266a0e4cb6a2fc633e57f84b1783f95c', 'skin_s3': '150412b1a9ff4489756567f970a9663b5c894fd50093241316f5e72ef7008e49', 'skin_s4': 'df44e89d5efbdbfb5d3702e92921ebc38060ee5df02bbb8a87604467780013ea', 'tonsil_s0': '8edd68ce0858c303386ed017264b89d7ffa90ba61dec9cc937ed849668394493', 'tonsil_s1': '3dc4adb307b85ac7e53bcf65a8d77ce02b62f058d58fcab8543618e804f815e9'}, '40x': {'breast_s0': 'f8ddf4696e0a0dd927a8b0c1e24242d3f4575190c7a26fe80f78d0932bb23227', 'breast_s1': '5fea33e2716a92054f02f3d8d75ca96a37f7e75acf91821451f6bd771d2fbc00', 'breast_s3': '457a3fdabaad46f8e007958419a5828a6d50e438f236b1bb0064b8d604712730', 'breast_s6': '64036da240ccf08273156939dbe1a3545951d07d1f62d97db7d3a3f7c8758689', 'cervix_s0': 'dc980ba27a43a6fe826e02cf529cf17e35f98a22781a530cd5c2b298ebb89d1c', 'colon_s1': '27a31ef22fe716285f63dab79bf45aa9cdf6e8e860b8ca725910a239b800050c', 'colon_s2': 'ad1d3e2cfa20c39da3c4db03c4760732a0e4d851b3ef7f8b45f30cb6ba74756f', 'heart_s0': 'cfa6690e8be28714583edb866cb3fadf575733cdd084b6d3e8cc68e49b3917d8', 'kidney_s0': '70ea019dc4c89647fe4ea34f65061b1386cac46cf69f78e036b1d3f14b6c1766', 'kidney_s1': '88fd5484af6f5d1734a7f0f7a8fe6e50151a39c73c48a11892baf79ffd1f724e', 'liver_s0': 'b6f74e1ecfaa59bbb75921ba6855877f5692487e6890e7a396cd0684d7cb467d', 'liver_s1': 'a6058b2e2ed771e7cf82dfc1f22c3590fe99d003abdd22d18cbe9266f44bc03d', 'lung_s1': '041fbcae0f18ae626c2b5daa37231020bad1edf589ea5149418ca9b727eae2c8', 'lung_s3': '4c2ad9337adb0683604f0c8683288cb8eb902f095dc09eeb8764c58ab6c75386', 'lymph_node_s0': 'e1b4659526ee64a22eada40cec267a7a1b75fc3079d84f6123bfe547bf2716b2', 'ovary_s0': '9e561f24a1cc9b6a3dfd9e2c2320c5d11a9ebae9d54bc5ddd6d9db3c446b5375', 'ovary_s1': '79f100d35cedee540d5f5319549a381faf8ebde736601e329a99f815b3978980', 'pancreatic_s0': 'f154d79606c8d1acff94c032bfc975e8518913f488da2bc9224d8ae32162c05b', 'pancreatic_s1': 'e46c59592eea273450012d599dc09e87a970dab35eae1e9f00f431ecdd1201ed', 'pancreatic_s2': 'efccb9598ad581bcc0da9fd7c4b2b585b7c405774f0359e2380dd2ae51f56a54', 'prostate_s0': '5397a52b30862f60780a8d105cacfa5a5ef7db66896edc88def2a6aa7b55cb3e', 'skin_s1': '4b4faf5248e8555fd68395f16d7a6ca653c02f7b45171da224f77026e5246f8c', 'skin_s2': '18b149fae149ac1b35c09701d8815b5818c90b5670643269a465d6aa34597b62', 'skin_s3': '2cd2e0c9d49c56192cb2eec65d2f4bec3feb0fef3a986472a2c50113e61e5bd5', 'skin_s4': 'e5879ec0348b02eb1c6cb8ca04b7c33f88d50328e95008309453ef24e21e1702', 'tonsil_s0': '0f95581fdc179715ac190df5e6a02df03e48baf156c4c700fc011101ba9315ff', 'tonsil_s1': '5f40a96f6c11cdfba8f80f4eb0bdf14c0b34000166c5d116a3714bf25e478222'}}
def get_sthelar_data( path: Union[os.PathLike, str], magnification: Literal['20x', '40x'] = '20x', slides: Optional[Sequence[str]] = None, download: bool = False) -> str:
366def get_sthelar_data(
367    path: Union[os.PathLike, str],
368    magnification: Literal["20x", "40x"] = "20x",
369    slides: Optional[Sequence[str]] = None,
370    download: bool = False,
371) -> str:
372    """Download the STHELAR dataset and convert its patches to hdf5 files.
373
374    NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use `slides` for a subset. The parquet files
375    of the slides are kept, since a parquet file can contain the patches of several slides.
376
377    Args:
378        path: Filepath to a folder where the data is downloaded for further processing.
379        magnification: The choice of magnification. Either '20x' or '40x'.
380        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
381        download: Whether to download the data if it is not present.
382
383    Returns:
384        Filepath where the converted patches are stored, in one sub-folder per slide.
385    """
386    slides = _validate(magnification, slides)
387    data_dir = os.path.join(path, magnification)
388    preprocessed_dir = os.path.join(data_dir, "preprocessed")
389
390    pending = [s for s in slides if not os.path.exists(os.path.join(preprocessed_dir, s, COMPLETE_MARKER))]
391    if not pending:
392        return preprocessed_dir
393
394    os.makedirs(os.path.join(data_dir, "data"), exist_ok=True)
395    os.makedirs(os.path.join(data_dir, "cell_metadata"), exist_ok=True)
396
397    lookups = {}
398    for slide in pending:
399        relative_path = f"cell_metadata/{slide}_cell_metadata.parquet"
400        metadata_path = os.path.join(data_dir, relative_path)
401        util.download_source(
402            path=metadata_path, url=_resolve_url(magnification, relative_path), download=download,
403            checksum=METADATA_CHECKSUMS[magnification][slide],
404        )
405        lookups[slide] = _load_lookup(metadata_path)
406        os.makedirs(os.path.join(preprocessed_dir, slide), exist_ok=True)
407
408    n_cpus = len(os.sched_getaffinity(0)) if hasattr(os, "sched_getaffinity") else (os.cpu_count() or 1)
409    n_workers = min(8, n_cpus)
410
411    shard_ids = sorted({i for slide in pending for i in SLIDE_SHARDS[magnification][slide]})
412    for shard_id in shard_ids:
413        relative_path = f"data/{_shard_name(magnification, shard_id)}"
414        shard_path = os.path.join(data_dir, relative_path)
415        util.download_source(
416            path=shard_path, url=_resolve_url(magnification, relative_path), download=download,
417            checksum=SHARD_CHECKSUMS[magnification][shard_id],
418        )
419        _convert_shard(shard_path, set(pending), lookups, preprocessed_dir, n_workers)
420
421    for slide in pending:
422        with open(os.path.join(preprocessed_dir, slide, COMPLETE_MARKER), "w"):
423            pass
424
425    return preprocessed_dir

Download the STHELAR dataset and convert its patches to hdf5 files.

NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use slides for a subset. The parquet files of the slides are kept, since a parquet file can contain the patches of several slides.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • magnification: The choice of magnification. Either '20x' or '40x'.
  • slides: The slides to use, see SLIDES for the valid choices. By default all slides are used.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the converted patches are stored, in one sub-folder per slide.

def get_sthelar_paths( path: Union[os.PathLike, str], magnification: Literal['20x', '40x'] = '20x', slides: Optional[Sequence[str]] = None, download: bool = False) -> List[str]:
428def get_sthelar_paths(
429    path: Union[os.PathLike, str],
430    magnification: Literal["20x", "40x"] = "20x",
431    slides: Optional[Sequence[str]] = None,
432    download: bool = False,
433) -> List[str]:
434    """Get paths to the STHELAR data.
435
436    Args:
437        path: Filepath to a folder where the data is downloaded for further processing.
438        magnification: The choice of magnification. Either '20x' or '40x'.
439        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
440        download: Whether to download the data if it is not present.
441
442    Returns:
443        List of filepaths for the hdf5 files, which contain the image ('image') and the labels
444        ('labels/instances' and 'labels/semantic') of a patch.
445    """
446    slides = _validate(magnification, slides)
447    preprocessed_dir = get_sthelar_data(path, magnification, slides, download)
448
449    data_paths = []
450    for slide in slides:
451        data_paths.extend(natsorted(glob(os.path.join(preprocessed_dir, slide, "*.h5"))))
452
453    assert len(data_paths) > 0
454    return data_paths

Get paths to the STHELAR data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • magnification: The choice of magnification. Either '20x' or '40x'.
  • slides: The slides to use, see SLIDES for the valid choices. By default all slides are used.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image ('image') and the labels ('labels/instances' and 'labels/semantic') of a patch.

def get_sthelar_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], magnification: Literal['20x', '40x'] = '20x', slides: Optional[Sequence[str]] = None, label_type: Literal['instances', 'semantic'] = 'instances', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
457def get_sthelar_dataset(
458    path: Union[os.PathLike, str],
459    patch_shape: Tuple[int, int],
460    magnification: Literal["20x", "40x"] = "20x",
461    slides: Optional[Sequence[str]] = None,
462    label_type: Literal["instances", "semantic"] = "instances",
463    resize_inputs: bool = False,
464    download: bool = False,
465    **kwargs
466) -> Dataset:
467    """Get the STHELAR dataset for nucleus segmentation and cell type classification.
468
469    Args:
470        path: Filepath to a folder where the data is downloaded for further processing.
471        patch_shape: The patch shape to use for training.
472        magnification: The choice of magnification. Either '20x' or '40x'.
473        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
474        label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
475            cell types of the nuclei, with the ids given by `CLASS_NAMES`.
476        resize_inputs: Whether to resize the inputs to the patch shape.
477        download: Whether to download the data if it is not present.
478        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
479
480    Returns:
481        The segmentation dataset.
482    """
483    if label_type not in ("instances", "semantic"):
484        raise ValueError(f"'{label_type}' is not a valid label type. Choose either 'instances' or 'semantic'.")
485
486    data_paths = get_sthelar_paths(path, magnification, slides, download)
487
488    if resize_inputs:
489        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True}
490        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
491            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
492        )
493
494    return torch_em.default_segmentation_dataset(
495        raw_paths=data_paths,
496        raw_key="image",
497        label_paths=data_paths,
498        label_key=f"labels/{label_type}",
499        patch_shape=patch_shape,
500        ndim=2,
501        with_channels=True,
502        **kwargs
503    )

Get the STHELAR dataset for nucleus segmentation and cell type classification.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • magnification: The choice of magnification. Either '20x' or '40x'.
  • slides: The slides to use, see SLIDES for the valid choices. By default all slides are used.
  • label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the cell types of the nuclei, with the ids given by CLASS_NAMES.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_sthelar_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], magnification: Literal['20x', '40x'] = '20x', slides: Optional[Sequence[str]] = None, label_type: Literal['instances', 'semantic'] = 'instances', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
506def get_sthelar_loader(
507    path: Union[os.PathLike, str],
508    batch_size: int,
509    patch_shape: Tuple[int, int],
510    magnification: Literal["20x", "40x"] = "20x",
511    slides: Optional[Sequence[str]] = None,
512    label_type: Literal["instances", "semantic"] = "instances",
513    resize_inputs: bool = False,
514    download: bool = False,
515    **kwargs
516) -> DataLoader:
517    """Get the STHELAR dataloader for nucleus segmentation and cell type classification.
518
519    Args:
520        path: Filepath to a folder where the data is downloaded for further processing.
521        batch_size: The batch size for training.
522        patch_shape: The patch shape to use for training.
523        magnification: The choice of magnification. Either '20x' or '40x'.
524        slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used.
525        label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
526            cell types of the nuclei, with the ids given by `CLASS_NAMES`.
527        resize_inputs: Whether to resize the inputs to the patch shape.
528        download: Whether to download the data if it is not present.
529        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
530
531    Returns:
532        The DataLoader.
533    """
534    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
535    dataset = get_sthelar_dataset(
536        path, patch_shape, magnification, slides, label_type, resize_inputs, download, **ds_kwargs
537    )
538    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the STHELAR dataloader for nucleus segmentation and cell type classification.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • magnification: The choice of magnification. Either '20x' or '40x'.
  • slides: The slides to use, see SLIDES for the valid choices. By default all slides are used.
  • label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the cell types of the nuclei, with the ids given by CLASS_NAMES.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.