Applied Research Engineer - Post-Training
London, England, United Kingdom · Tempo total
Seja o primeiro a se candidatar
- Experiência
- Qualquer
- Salário
- —
- Vagas
- 1
- Publicado
- há 2 horas
- Modo de trabalho
- No escritório
- Educação
- MSc or PhD in Machine Learning, Computational Biology, or related fields or equivalent experience
- Retomar
- Obrigatório candidatar-se
Onde você trabalhará
Descrição da vaga
About Helical
Helical is revolutionizing biological research by creating in-silico laboratories for drug discovery. Traditional wet lab methods remain slow, costly, and limited by trial-and-error processes. Helical aims to accelerate this by providing an application platform that leverages biological foundation models, allowing pharmaceutical and biotech organizations to conduct millions of virtual experiments rapidly, significantly reducing timelines from years to days. As a fast-growing, founder-led company based in Europe with a deep commitment to quality and ownership, Helical offers a dynamic environment where complexity and responsibility are embraced.
Role Overview
The Applied Research Engineer - Post-Training will be responsible for managing the entire post-training cycle of biological foundation models. This includes developing strategies to align these general-purpose models with therapeutic objectives, creating pipelines to specialize these models for pharmaceutical clients, and applying model adaptations tailored to specific disease areas, cell types, and perturbation contexts critical to drug discovery phases such as target identification and hit discovery.
Key Responsibilities
- Design and deploy post-training pipelines that customize biological foundation models for precise therapeutic scenarios and client requirements.
- Develop validation systems that tie improvements in models to biological realities, utilizing embeddings, perturbation datasets, and external tools such as OpenTargets.
- Manage the full experimental lifecycle, from hypothesis formulation, running distributed GPU-based training sessions, through to detailed analysis and delivery to clients.
- Collaborate closely with machine learning infrastructure engineers and biologists to ensure scientific validity and effective deployment of models.
- Contribute to open-source projects like the helical-package and influence the strategic development of post-training capabilities as the company scales.
- Stay updated with the forefront of post-training research and integrate applicable advancements into the production environment.
Qualifications and Skills
- Master's degree or PhD in Machine Learning, Computational Biology, or related disciplines, or significant equivalent professional experience.
- Practical experience with post-training strategies such as fine-tuning, LoRA, DPO, RLHF, or related alignment techniques.
- Advanced proficiency in Python and PyTorch, including writing training loops, debugging distributed computational runs, and hands-on work with model internals.
- Well-versed in transformer neural network architectures and their practical applications.
- Proven ability to design, execute, and evaluate experiments with strong methodological rigor, tracking performance metrics and iterating on approaches with sound analytical insights.
- Highly autonomous work ethic with strong decision-making capacity despite limited data and resources, owning problems from start to finish.
- Excellent communication skills to articulate technical considerations clearly to interdisciplinary teams including ML specialists, biologists, and product managers.
Preferred Extras
- Experience with biological foundation models such as Geneformer, scGPT, ESM, or general computational biology expertise.
- Understanding of drug discovery pipelines, including target identification and perturbation biology concepts.
- Demonstrated success in deploying post-training enhancements into operational production platforms.
- Familiarity with distributed training environments involving multi-GPU, multi-node setups, and frameworks like NCCL, DeepSpeed, or FSDP.
- Published research in prominent ML or computational biology conferences and journals (NeurIPS, ICML, ICLR, Nature Methods, etc.).
- Contributions to open-source machine learning tools or software.