Single-cell sequencing may be more specific than bulk sequencing but comes with a significant cost. At FoG Live: Single Cell and Spatial, we were joined by Jun Ding (McGill University) who discussed his lab’s solution to this problem: scSemiProfiler. This computational tool integrates deep generative AI and active learning to ‘semi-profile’ single-cell data for any studied cohort, based on bulk data and single-cell templates from a few representative samples.
Development of Single Cell Semi-Profiling
scSemiProfiler was developed in response to concerns about the affordability of single cell sequencing, despite its usefulness. While bulk sequencing can be performed for a fraction of the cost, insights from this method are limited in comparison to single-cell sequencing, giving only the average gene expression for a group of cells.
The single-cell semi-profiler is a pipeline designed to make the insights from single-cell sequencing more widely accessible. The pipeline involves initial bulk sequencing followed by clustering. Then, representative samples from each cluster are used for single-cell sequencing. Using a combination of this data, single-cell insights can be inferred for all samples in a cohort.
Active learning and VAE-GAN
To obtain high quality results, you need to use the best representative samples available. This is where an active learning algorithm comes in. After representative samples are taken from the initial clusters, the most heterogenous groups can be split further for better accuracy. Using these new clusters, better representative samples can be chosen for further analysis.
Semi-profiling utilises a deep learning model, a variational encoder known as VAE-GAN. This model takes representative single-cell data and generates reconstructed data for the other cells. The model is trained so that a ‘discriminator’ cannot tell the difference between the real- and semi-profiled cells, ensuring they are as true to life as possible.
However, the representative cells, and therefore the inferred data, will not precisely mirror the true data in the first instance. As such, the ‘template’ of inferred data needs to be ‘pushed’ towards the target in a reconstruction phase. The initial bulk sequencing data is used to inform the direction of this movement, bringing the data closer to real life. In experimental samples, after different training stages, the template of the inferred cells gradually shifts from the representative sample towards the ‘true’ target samples.
What next?
After single-cell data (real-profiled or inferred using semi-profiling) has been generated, it can be used for downstream analysis. This includes biomarker discovery, cell-cell interaction and pseudotime analysis, deconvolution, visualisation and a range of other applications that typically use traditional single-cell sequencing data.
Furthermore, in one example, Jun Ding described a study in a COVID-19 patient cohort. The correlation between the generated and real single-cell data was high, with over 0.9 correlation in a 124-person cohort with 28 representative samples. In addition, subsequent biomarker discovery, cell-cell interaction and pseudotime analysis using the two different groups showed very similar results.
These results show that semi-profiling is an efficient method and a cost-effective alternative to single-cell sequencing. Jun Ding posits that researchers using semi-profiling can save 30-50% of the cost of performing single-cell sequencing on every sample.
References and further reading
Wang, J., Fonseca, G.J. & Ding, J. scSemiProfiler: Advancing large-scale single-cell studies through semi-profiling with deep generative models and active learning. Nat Commun 15, 5989 (2024). https://doi.org/10.1038/s41467-024-50150-1.


