- خبرة
- أكثر من 5 سنوات
- مرتب
- —
- الوظائف الشاغرة
- 1
- تم النشر
- • 4 قطع
- وضع العمل
- العمل من المنزل
- سيرة ذاتية
- مطلوب للتقديم
المسمى الوظيفي
About distil labs
distil labs develops a platform that automatically fine-tunes task-specific small language models. This solution enables customers to replace general-purpose large language models (LLMs) in their agent systems with smaller, purpose-built alternatives that maintain quality on specific tasks while reducing costs by 50-80% and lowering latency. The platform employs production data traces collected from customers to create synthetic training data, train compact models that rival state-of-the-art performance on narrow tasks, and deploy these models to OpenAI-compatible endpoints hosted in the cloud, on-premises, or at the edge.
Role Overview
As a Machine Learning Engineer at distil labs, you will be responsible for managing the complete lifecycle of every model delivered—spanning synthetic data generation, distributed training and evaluation, and deployment to production environments. Collaborating closely with the science team, you will help develop innovative techniques in knowledge distillation, synthetic data creation, and model self-improvement. Your work will translate these methods into scalable production pipelines that elevate the quality of models for all customers. Being part of a small team, your contributions will significantly influence technical decisions, product strategy, and company culture.
Key Responsibilities
- Manage the end-to-end model lifecycle: starting from raw customer data traces through synthetic data generation, validation, distributed fine-tuning, evaluation, and ultimately deploying models to production endpoints.
- Partner with the research team to pioneer new approaches in knowledge distillation, synthetic data generation, and self-improving models, implementing the successful ideas into automated production pipeline components.
- Design and implement evaluation and benchmarking tools that rigorously quantify model improvements and run experiments to validate changes.
- Operate and optimize distributed fine-tuning workloads leveraging technologies including HuggingFace, PyTorch, DDP/FSDP, and LoRA on cloud and on-premises GPU infrastructure.
- Manage the computing infrastructure supporting these pipelines, including Argo Workflows, Kubernetes orchestration, and maintaining a secure, multitenant, low-latency serving layer using vLLM and FastAPI under heavy load conditions.
Candidate Qualifications
- Minimum 5 years of experience deploying machine learning systems to production, with direct responsibility for model training, evaluation, and deployment workflows.
- Strong ability to interpret and implement machine learning research papers, collaborating effectively with scientists on methodological development beyond just technical execution.
- Meticulous experimental design skills to objectively evaluate model changes through quantitative measurements rather than subjective intuition.
- Expert-level proficiency in Python programming specifically tailored to machine learning, comfortable using HuggingFace and PyTorch frameworks.
- Experienced in running distributed training jobs with frameworks such as Distributed Data Parallel (DDP), Fully Sharded Data Parallel (FSDP), DeepSpeed, Ray, or Kubeflow, including troubleshooting common failure scenarios.
- Practical expertise in managing Kubernetes clusters, Argo Workflow orchestration, container ecosystems, and cloud platforms such as AWS, GCP, or Azure in production environments.
Additional Desirable Skills
- Hands-on experience with knowledge distillation, synthetic data generation, or model compression techniques.
- The ability to assess training data quality critically, identifying synthetic examples that may degrade model performance.
- Knowledge of inference optimization methods including quantization, sparsity, compilation, and familiarity with serving engines like Triton, TensorRT-LLM, or vLLM.
- Proficiency in infrastructure-as-code tools (e.g., Terraform, Pulumi), observability platforms (such as Prometheus, Grafana, Datadog), and GPU resource cost optimization.
- Experience managing hybrid or on-premises GPU clusters and high-performance storage systems (NVMe, Infiniband).
- Background contributing to academic research with publications in machine learning or related disciplines and participation in open-source ML infrastructure projects.
Why Join distil labs?
- Full ownership over the entire model lifecycle with autonomy to select the best tools.
- A close connection between research and production, enabling new scientific methods to impact customers quickly.
- An opportunity to enable advanced AI solutions for organizations lacking extensive GPU resources or machine learning expertise, powering clients in defense, cybersecurity, education technology, and robotics.
- Competitive salary with equity options and a meaningful role in guiding company growth.
- A remote-first European focus allowing work from anywhere within EU time zones, complemented by in-person gatherings in Berlin.
Equal Opportunity Statement
distil labs is committed to fostering diversity and inclusion, providing equal employment opportunities regardless of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability.