|
Persistent Identifier
|
perma:LIST.FBELJU |
|
Publication Date
|
2026-07-06 |
|
Title
| GigaSOM.jl: High-performance clustering and visualization of huge cytometry datasets [* Cross-Reference *] |
|
Other Identifier
| https://doi.org/10.1093/gigascience/giaa127
OpenAlex ID: https://openalex.org/W3099773610 |
|
Author
| Miroslav Kratochvíl (Charles University, Czech Academy of Sciences, Institute of Organic Chemistry and Biochemistry) - ORCID: https://orcid.org/0000-0001-7356-4075
Oliver Hunewald (Luxembourg Institute of Health) - ORCID: https://orcid.org/0000-0001-5402-5084
Laurent Heirendt (University of Luxembourg) - ORCID: https://orcid.org/0000-0003-1861-0037
Vasco Verissimo (University of Luxembourg) - ORCID: https://orcid.org/0000-0003-3884-9125
Jiřı́ Vondrášek (Czech Academy of Sciences, Institute of Organic Chemistry and Biochemistry) - ORCID: https://orcid.org/0000-0002-6066-973X
Venkata Satagopam (University of Luxembourg, Luxembourg Institute of Science and Technology) - ORCID: https://orcid.org/0000-0002-6532-5880
Reinhard Schneider (University of Luxembourg, Luxembourg Institute of Science and Technology) - ORCID: https://orcid.org/0000-0002-8278-1618
Christophe Trefois (University of Luxembourg, Luxembourg Institute of Science and Technology) - ORCID: https://orcid.org/0000-0002-8991-6810
Markus Ollert (University of Southern Denmark, Odense University Hospital, Luxembourg Institute of Health) - ORCID: https://orcid.org/0000-0002-8055-0103 |
|
Point of Contact
|
Use email button above to contact.
LIST QDKM (LIST) |
|
Description
| Background: The amount of data generated in large clinical and phenotyping studies that use single-cell cytometry is constantly growing. Recent technological advances allow the easy generation of data with hundreds of millions of single-cell data points with >40 parameters, originating from thousands of individual samples. The analysis of that amount of high-dimensional data becomes demanding in both hardware and software of high-performance computational resources. Current software tools often do not scale to the datasets of such size; users are thus forced to downsample the data to bearable sizes, in turn losing accuracy and ability to detect many underlying complex phenomena. Results: We present GigaSOM.jl, a fast and scalable implementation of clustering and dimensionality reduction for flow and mass cytometry data. The implementation of GigaSOM.jl in the high-level and high-performance programming language Julia makes it accessible to the scientific community and allows for efficient handling and processing of datasets with billions of data points using distributed computing infrastructures. We describe the design of GigaSOM.jl, measure its performance and horizontal scaling capability, and showcase the functionality on a large dataset from a recent study. Conclusions: GigaSOM.jl facilitates the use of commonly available high-performance computing resources to process the largest available datasets within minutes, while producing results of the same quality as the current state-of-art software. Measurements indicate that the performance scales to much larger datasets. The example use on the data from a massive mouse phenotyping effort confirms the applicability of GigaSOM.jl to huge-scale studies. (2020-11-01)
***This entry has been automatically imported via OpenAlex by LIST harvest scripts. Please refer to https://doi.org/10.1093/gigascience/giaa127 for the original and latest version of the publication*** (2026-07-01) |
|
Subject
| Astronomy and Astrophysics; Chemistry; Medicine, Health and Life Sciences; Physics |
|
Keyword
| Visualization
Computer science
Cluster analysis
Information retrieval
Data mining
Data science
Artificial intelligence |
|
Topic Classification
| Single-cell and spatial transcriptomics
Cell Image Analysis Techniques
Digital Holography and Microscopy |
|
Deposit Date
| 2020-11-01 |
|
Data Type
| Article |
|
Data Source
| GigaScience |