Clinical genomics
WGS, WES and targeted panels: alignment, post-processing, variant calling with multiple callers, annotation and automatic variant classification. Pharmacogenomic (PGx) analysis and computation of polygenic and clinical risk scores.
Senior Bioinformatics Scientist
> ▌
I build pipelines and analysis tools for clinical and multi-omics data: genomics, transcriptomics, proteomics, microbiome. Most of my work is on NGS data and on bringing heterogeneous sources together, from omics layers to real-world data, into digital twins: virtual replicas of a person's health status, at the service of cancer research and precision medicine.
A few words about me.
I'm a Senior Bioinformatics Scientist with a background in Computer Science and a Ph.D. in Bioinformatics. Since 2018 I have been designing algorithms and pipelines for clinical and multi-omics datasets, with particular experience on NGS data (short and long reads) across genomic, transcriptomic, proteomic and microbiomic layers.
Most of my work goes into building and optimising analysis workflows, from preprocessing to downstream statistical and integrative analyses. I use R and Python daily, manage workflows with Nextflow and Docker, and reach for machine learning or network modelling when they genuinely help.
Almost every project I work on starts by combining omics, clinical and real-world data to find molecular patterns that are useful in the clinic. It only works as a team effort: much of the result depends on talking to the people who produce the data and to the people who will use it.
From the sequencer output to biological and clinical interpretation.
WGS, WES and targeted panels: alignment, post-processing, variant calling with multiple callers, annotation and automatic variant classification. Pharmacogenomic (PGx) analysis and computation of polygenic and clinical risk scores.
Bulk RNA-seq (QC, splice-aware alignment, quantification, differential expression, functional enrichment) and single-cell work (cell-level QC, integration, clustering, annotation, trajectories, cell-cell communication). miRNA profiling and miRNA/mRNA networks.
16S and shotgun metagenomics pipelines, from raw reads to taxonomic and functional profiles. Clinical-style indicators such as the Firmicutes/Bacteroidetes ratio, alpha and beta diversity, enterotypes, dysbiosis and pathogen detection.
Personalised digital models fed by omics, clinical and real-world data, wearables included. Harmonisation of the sources, standardised indicators for each layer, and simulation of disease trajectories and treatment response.
LC/MS pipelines from raw files to the protein by sample matrix: run and batch QC, identification with FDR control, label-free, TMT or DIA quantification, protein inference, normalisation, missing-value handling and statistical testing.
Methylation from NGS (WGBS, RRBS) with CpG-level extraction and DMC/DMR discovery; EPIC v2 arrays with normalisation, cell-type deconvolution and epigenetic clocks for biological age.
Biomedical knowledge graphs in Neo4j (pharmacogenomics, microbiome, supplements), GraphRAG and rule-based reasoning engines, NER for automatic graph expansion, evidence-based disease scoring and machine learning for molecular signatures.
Reproducible, containerised workflows, cloud deployment, versioning and data traceability. Data schemas and ETL-style steps (ingest, validate, standardise, store, publish) with automated QC checkpoints.
Predictive modelling, survival analysis, dimensionality reduction, feature selection, ROC curves and classification performance assessment, with attention to model interpretability.
Four things I care about when setting up an analysis.
It starts with quality control and ends with numbers someone can actually act on. Containerised workflows that give the same result on every run.
A single layer only tells half the story. I harmonise datasets that don't naturally match and make them talk to each other through statistics, interaction networks and knowledge graphs.
Explicit definitions, confounders under control (depth, batch, protocol) and proper references when a result is meant to be clinical. An explainable indicator beats an opaque predictor.
I work shoulder to shoulder with biologists, clinicians, data scientists and software engineers. A good part of the value lies in making a complex output readable.
A selection of recent work.
A virtual replica of a person's health status, fed by wearables, blood tests, phenotypic data, genomics and microbiome. Each layer is harmonised and summarised into standardised indicators that flow into the platform.
A knowledge graph joining pharmacogenomics, microbiome and supplements on shared entities such as disease and drug. The score accumulates evidence from microbiome, genomics and phenotypes, and shows the graph paths that justify it.
A patented platform that starts from FASTQ, BAM or VCF and produces a full clinical report, through QC, alignment, variant calling and annotation. Variants are interpreted against the main oncology, pharmacology and regulatory databases.
Differential diagnosis between prostate cancer and benign hyperplasia in a cohort with PSA between 2.5 and 10 ng/mL, where PSA alone says little. Counting single vesicles under super-resolution microscopy, the STAT3, CyclinD1 and CD81 combination reaches an AUC of about 0.77 against 0.67 for PSA.
How microgravity affects gene expression in a 3D bone-like model, within an experiment carried out in orbit. Integrated RNA-seq and miRNA-seq analysis points to modules centred on Wnt5b and Runx2, consistent with a rewiring of osteoblast differentiation circuits.
Crosstalk between miRNAs, mRNAs and the proteome in PD-1 positive NK cells. Network analysis shows that modulation of certain miRNAs matches inverse changes in mRNA and protein levels, a hint of post-transcriptional control over checkpoints.
A framework linking colorectal cancer associated microbial shifts to the metabolic tendencies of the community. A score balances the abundance of producer and consumer microbes for each metabolite and returns a ranked list.
Doctoral work on models and algorithms for graph motifs, the small patterns that recur in interaction networks. The model estimates the expected number of motifs without enumerating every possible configuration, cutting computation time considerably.
More than 30 peer-reviewed works across computational oncology, multi-omics and methods.
Full list (30+ papers) on Google Scholar ↗
For scientific collaborations or data analysis projects, drop me a line.