Machine Learning Engineer - Inference Optimization
Remote · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 6 hours ago
- Work mode
- Work from home
- Resume
- Required to apply
Job description
Overview
This opportunity is offered through a partner organization looking for a Machine Learning Engineer focusing on inference optimization based in Ireland. The role centers on enhancing the performance of advanced ML systems deployed in production settings.
This position involves bridging research and engineering by transforming innovative models into swift, dependable, and cost-effective AI solutions. Your efforts will influence scalability, user experience, and the operational efficiency of AI-driven products.
You will engage deeply with optimization efforts from model structure, GPU execution, to managing large-scale inference frameworks. Collaborating with research, infrastructure, and product teams, you will enable advanced AI capabilities.
Key Responsibilities
- Enhance inference systems to boost latency, throughput, scalability, and reduce operational expenses.
- Analyze and profile GPU and CPU inference pipelines to detect bottlenecks in components like memory handling, kernel operations, data batching, and flow.
- Apply sophisticated optimization techniques such as quantization, key-value caching improvements, speculative decoding, batching, streaming, and model simplification.
- Work alongside research engineers to transition experimental model architectures into dependable production systems.
- Develop and maintain inference-serving infrastructure with modern frameworks, custom runtimes, or specialized serving platforms.
- Evaluate and benchmark performance across diverse hardware, including GPUs, CPUs, and cloud environments.
- Enhance system dependability, monitoring, observability, and minimize costs during real production workloads.
- Participate in refining engineering practices to enhance quality, scalability, and maintainability of ML infrastructure.
Qualifications and Experience
- Substantial experience in machine learning inference optimization or building high-performance ML systems.
- In-depth knowledge of ML principles including neural network designs, attention mechanisms, memory management, and compute graph theory.
- Expertise in deep learning frameworks like PyTorch and deploying models in production contexts.
- Strong background in GPU performance tuning with technologies such as CUDA, ROCm, Triton, or kernel-level optimization.
- Proven track record of scaling inference setups beyond research prototypes for actual user scenarios.
- Proficiency in programming spanning machine learning and systems engineering.
- Ability to work autonomously in dynamic environments with shifting priorities.
- Familiarity with inference platforms like TensorRT, ONNX Runtime, vLLM, or Triton is advantageous.
- Knowledge of large language models, long-context inference, distributed architectures, low-latency services, or hardware optimization is a plus.
- Contributions to open-source ML tools or inference-related projects are considered beneficial.
Benefits
- Competitive salary complemented by significant equity shares.
- Chance to work on performance-critical AI systems influencing real products.
- High ownership level over infrastructure impacting scalability and operational efficiency.
- Close collaboration with cutting-edge research, infrastructure, and product teams.
- Engagement with advanced ML technologies and practical AI applications.
- Culture centered on engineering excellence, experimentation, and quality assurance.
- Flexible arrangements supporting remote work.
- Opportunity to contribute to an innovative AI-driven organization’s expansion.