Machine Learning Scientist / Computational Biologist
Singapore · Full Time
Be the first to apply
- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 7時間前
- Work mode
- In office
- Education
- PhD or MSc in computational biology or related field
- Resume
- Required to apply
Where you'll work
Job description
About the Role
Join an early-stage, venture-backed biotechnology company focused on advancing antibody-based therapeutics in immunology. The position serves as the computational ally to the R&D team, primarily analyzing complex biological datasets to evaluate the potential of targets and molecules. These datasets include transcriptomics, genomics, proteomics, sequences, campaign results, and assay data. The role also involves developing pipelines and data frameworks to ensure rapid, reproducible, and consistent analysis throughout the discovery program.
Key Responsibilities
- End-to-end transcriptomic data analysis including single-cell RNA-seq, covering quality control, data integration, batch correction, clustering, cell-type identification, and differential gene expression.
- Analyze and combine various omics data such as genomics, proteomics, and epigenomics from bulk and single-cell experiments, and resolve conflicting findings.
- Evaluate target expression profiles and biology across tissues, cell types, and disease versus healthy states to inform target validation, indication choice, and patient stratification hypotheses.
- Integrate both public and proprietary omics, biomarker, and human genetic data to form a robust evidence base supporting target decisions.
- Transform data analyses into clear, actionable conclusions about target viability and conditions that would invalidate those conclusions.
- Conduct antibody sequence analyses including numbering, germline assignment, CDR annotation, and clonal diversity assessment.
- Interpret next-generation sequencing data from display and repertoire campaigns to guide selection and decision-making.
- Assess target sequences for conservation, cross-species homology relevant to animal models, isoform diversity, and variant landscapes.
- Analyze laboratory assay and screening data from both internal and partner sources, ensuring quality control, normalization, and comparative analyses.
- Develop and maintain computational pipelines converting raw instrument and sequencing outputs into analysis-ready datasets.
- Create a comprehensive R&D data model linking targets, molecular designs, sequences, and assay results to maintain clear data associations.
- Develop automated data ingestion from laboratory instruments, electronic lab notebooks, and external collaborators including CROs and cloud labs, setting data return criteria as needed.
- Maintain data integrity and provenance to enable result traceability and reproducibility.
- Design dashboards and reporting tools so teams can monitor design-to-result metrics without the need for custom analyses.
- Participate in in-silico antibody design by performing structure-based analysis of antibodies and antigens, supporting epitope selection and design hypotheses.
- Utilize computational methods to prioritize candidate binders and engineered variants before laboratory testing.
- Identify developability risks from sequence and structural data for engineering assessment.
- Implement machine learning models responsibly to support ranking, outcome prediction, and pattern recognition, ensuring performance is rigorously benchmarked against simpler methods.
- Provide clear, specific recommendations regarding experimental directions and resource allocation based on computational findings.
- Keep computational predictions and empirical data distinctly documented in all outputs.
- Evaluate external software tools and platforms, offering justified build-versus-buy decisions.
- Effectively communicate results to bench scientists, program leads, and senior leadership to facilitate informed decision-making.
Required Qualifications
- PhD with at least 3 years or MSc with 6 years of relevant experience in computational biology, bioinformatics, or a related quantitative field.
- Proven hands-on experience analyzing transcriptomic data including single-cell RNA-seq: quality control, integration, clustering, annotation, and differential expression, with capability to identify artefacts versus genuine biological signals.
- Experience across multiple omics types such as genomics, proteomics, or epigenomics, and understanding of their generation, normalization, and limitations.
- Comprehensive biological data analysis aptitude from raw data processing through quality assurance to actionable conclusions.
- Track record of applying machine learning to biological or omics data with expertise in feature engineering, model selection, evaluation, and avoidance of common pitfalls like data leakage and overfitting.
- Sound judgment on when machine learning is appropriate, favoring simpler models where justified.
- Direct experience working within drug discovery pipelines and managing supporting computational workflows.
- Advanced proficiency in Python and machine learning frameworks (e.g., scikit-learn, PyTorch), along with strong SQL skills and relational data modeling knowledge.
- Familiarity with reproducible engineering practices including version control, pipeline documentation, environment management, and provenance tracking.
- Strong statistical understanding, particularly in handling small-sample datasets, batch effects, and uncertainty quantification.
- Working knowledge of antibody biology and related experimental methodologies.
- Ability to collaborate directly with laboratory scientists and communicate data limitations clearly.
Preferred Qualifications
- Experience in in-silico protein or antibody design such as structure prediction, epitope analysis, sequence design, or developability prediction.
- Advanced deep learning experience with biological data, including protein language models or generative and structure-based approaches.
- Familiarity with antibody-specific analysis tools and numbering schemes, plus repertoire or display sequencing analysis.
- Experience integrating large-scale public omics atlases and multi-omics data.
- Expertise in proteomics or mass spectrometry data and human genetics-based variant interpretation.
- Use of electronic lab notebooks or laboratory information management systems and integration of experimental data from lab instruments.
- Experience managing data from contract research organizations and automated cloud labs.
- Knowledge of immunology, autoimmune and inflammatory diseases, oncology, cytokine biology, or receptor biology.
- Experience with workflow orchestration, cloud computational resources, and GPU-based workloads.
- Background working in early-stage biotech or venture-backed companies.
Role Responsibilities Summary
You will lead the computational foundation for target and program decisions through comprehensive omics and translational data analysis, antibody sequence bioinformatics, and the construction of scalable discovery pipelines that ensure reproducibility and accessibility of results. Your insights will drive what targets and projects are pursued or discontinued.
Ideal Candidate Backgrounds
Applicants with experience from computational biology or bioinformatics divisions within biotech or pharmaceutical companies, translational biology data science teams, multi-omics and single-cell genomics groups focused on therapeutic discovery, or academia specializing in computational immunology, functional genomics, or single-cell biology combined with robust software development and industry collaborations are encouraged to apply.