Bioinformatics Β· Cancer Genomics Β· Healthcare Data Engineering Β· Scientific Machine Learning
I am a Research Scientist, Bioinformatician, and Data Scientist working at the intersection of computational biology, cancer genomics, artificial intelligence, biostatistics, and precision medicine.
My research combines genomics, transcriptomics, multi-omics integration, machine learning, and large language models to transform complex biological and health data into evidence for:
- Biomarker discovery
- Therapeutic-target prioritization
- Precision oncology
- AI-assisted drug discovery
- Reproducible biomedical research
I currently contribute to cancer-genomics research as a Bioinformatician at the University of Arkansas for Medical Sciences (UAMS), developing scalable NGS workflows and analyzing TCGA and CPTAC datasets. I am also pursuing an MSc in Bioinformatics at Northeastern University, strengthening my expertise in translational bioinformatics, machine learning, and trustworthy biomedical AI.
Beyond biomedical research, I independently reproduced and extended machine-learning workflows for streamflow prediction and National Water Model bias correction. My HYDRO-FLOW-AI project builds on open workflows from the Alabama Water Instituteβs NWM-ML project and research involving a University of West Florida researcher, while remaining an independent project with no claim of institutional affiliation.
My long-term goal is to build trustworthy computational systems that connect biological evidence, clinical data, and artificial intelligence to accelerate scientific discovery and improve patient outcomes.
- 𧬠Computational biology and bioinformatics
- ποΈ Cancer genomics and precision oncology
- π§ͺ NGS analysis and variant interpretation
- π Multi-omics integration
- π Computational drug discovery
- π€ Artificial intelligence and deep learning
- π Biomedical large language models
- β‘ Agentic AI and retrieval-augmented generation
- βοΈ Cloud-based scientific computing
- π₯ Healthcare data engineering
- π Scientific machine learning for hydrology
| Project | Research focus | Repository |
|---|---|---|
| RTK/NRTK TNBC | Patient-level kinase alterations and drug-target prioritization in triple-negative breast cancer | View project |
| HYDRO-FLOW-AI | Extreme-aware streamflow prediction and National Water Model bias correction | View project |
| NIH Clinical Trials Lakehouse | Reproducible healthcare data engineering for clinical-trial analytics | View project |
| CDC Healthcare Streaming ETL | Streaming ingestion, validation, transformation, and public-health analytics | View project |
| TCGAβCPTAC Kafka Platform | Event-driven processing of large-scale cancer multi-omics data | View project |
| Synthetic Variant Calling Benchmark | Reproducible benchmarking of NGS variant-calling workflows | View project |
| Genomic Foundation Models | Transformer-based representation learning for genomic sequences | View project |
| USAG1 Validation | Computational evidence synthesis and therapeutic-target validation | View project |
HYDRO-FLOW-AI is an independent, extreme-aware machine-learning framework for streamflow prediction and site-specific National Water Model bias correction at USGS gauges.
- Leakage-safe temporal training, validation, and testing
- Historical USGS streamflow and climate-data integration
- Site-specific model-performance diagnostics
- Evaluation using RMSE, MAE, bias, and NSE
- Q95 and Q99 high-flow evaluation
- Peak-magnitude error analysis
- Extreme-event detection and threshold-based assessment
- XGBoost residual bias correction
- Quantile-regression uncertainty intervals
- LSTM and Transformer-based forecasting
- River-network graph neural networks
- Explainability and model-drift monitoring
Computational framework for rational drug-combination prioritization using complementary biological and pharmacological evidence.
Identification of compensatory kinase alteration patterns and potential drug-target combinations using TCGA cancer-genomics data.
LLM-supported retrieval, evaluation, and synthesis of biomedical evidence for research decision support.
Transformer-based representation-learning approaches for genomic sequences and downstream biological prediction.
Reproducible lakehouse and streaming architectures for clinical-trial, public-health, and biomedical data.
- AI for precision oncology
- Computational drug discovery
- Cancer multi-omics
- Biomedical large language models
- Genomic foundation models
- Agentic AI for scientific discovery
- Trustworthy and explainable AI
- Scalable scientific computing
- Extreme-aware streamflow forecasting
- MSc Bioinformatics β Northeastern University (in progress)
- MSc Molecular Biology (Bioinformatics) β UmeΓ₯ University
- BSc Biotechnology and Genetic Engineering β Khulna University
- Advanced Diploma in Data Science and Data Engineering
- Graduate Certificate in Project Management
- Health Informatics β Johns Hopkins University
- Business Analytics and Data-Driven Decision-Making β University of Toronto
- Project Management
- Cloud Computing
- Machine Learning and Data Science
- Trustworthy and causal AI
- Agentic and multi-agent systems
- Biomedical large language models
- Genomic foundation models
- Scalable multi-omics analytics
I welcome research and open-source collaboration in:
- Computational biology and bioinformatics
- Cancer genomics and precision medicine
- Biomedical artificial intelligence
- Computational drug discovery
- Healthcare data engineering
- Scientific machine learning
- Reproducible research software

