Machine Learning Engineer - Inference Optimization
Remote · ਪੂਰਾ ਸਮਾਂ
ਅਰਜ਼ੀ ਦੇਣ ਵਾਲੇ ਪਹਿਲੇ ਵਿਅਕਤੀ ਬਣੋ
- ਅਨੁਭਵ
- ਕੋਈ ਵੀ
- ਤਨਖਾਹ
- —
- ਖੁੱਲ੍ਹਣ ਵਾਲੀਆਂ ਥਾਵਾਂ
- 1
- ਪੋਸਟ ਕੀਤਾ ਗਿਆ
- 4 ਘੰਟੇ
- ਕੰਮ ਮੋਡ
- ਘਰੋਂ ਕੰਮ ਕਰੋ
- ਰੈਜ਼ਿਊਮੇ
- ਅਰਜ਼ੀ ਦੇਣ ਲਈ ਲੋੜੀਂਦਾ ਹੈ
ਕੰਮ ਦਾ ਵੇਰਵਾ
Overview
This opportunity is offered by a partner company based in Germany seeking a Machine Learning Engineer specialized in inference optimization. This role focuses on enhancing the efficiency and performance of advanced machine learning models deployed in production environments, balancing cutting-edge research and engineering practices to deliver scalable, cost-effective AI solutions.
Key Responsibilities
- Improve ML inference systems to optimize latency, throughput, scalability, and reduce operational costs.
- Analyze and profile GPU and CPU inference pipelines, addressing bottlenecks in memory, kernels, batching, and data flow.
- Apply advanced methods such as quantization, KV-cache optimization, speculative decoding, batching, streaming, and model simplification.
- Work closely with research engineers to transition novel model architectures into dependable production-ready systems.
- Develop and maintain inference-serving infrastructure utilizing modern frameworks, custom runtimes, or dedicated serving platforms.
- Conduct benchmarking of model performance across diverse hardware setups including GPUs, CPUs, and cloud environments.
- Enhance system reliability, monitoring, and cost efficiency to handle real-world production loads effectively.
- Drive engineering best practices to bolster scalability, maintainability, and quality of ML infrastructure.
Required Qualifications
- Demonstrated professional experience in optimizing ML inference and managing high-performance ML systems in production settings.
- Comprehensive understanding of machine learning fundamentals such as neural networks, attention mechanisms, memory optimization, and computational graphs.
- Hands-on proficiency with PyTorch or equivalent deep learning frameworks and deploying models operationally.
- Experience optimizing GPU performance using CUDA, ROCm, Triton, or similar technologies at the kernel level.
- Proven ability to scale inference systems beyond experimental or research environments to serve real users.
- Strong programming expertise bridging machine learning and systems engineering disciplines.
- Capability to manage fast-paced work environments with autonomy and accountability.
- Familiarity with inference platforms like TensorRT, ONNX Runtime, vLLM, or Triton is advantageous.
- Knowledge of large language models, extended context inference, distributed systems, low-latency services, or hardware-level optimization is a plus.
- Contributions to open-source ML or inference tools considered beneficial.
Benefits
- Attractive compensation including meaningful equity opportunities.
- Engagement with AI systems critical to product performance and impact.
- Significant ownership of infrastructure that influences system scalability and operational efficiency.
- Collaboration across research, infrastructure, and product teams.
- Exposure to advanced machine learning technologies and practical AI applications.
- A culture prioritizing engineering excellence, experimentation, and quality.
- Flexibility of a remote working environment.
- Chance to contribute to a forward-thinking AI-driven organization.
Additional Information
This position is managed by a partner company which handles applications and subsequent hiring stages. The recruitment workflow involves AI-driven matching to identify top candidates, but all final hiring decisions are human-made. The company ensures compliance with data privacy regulations including GDPR, and applicants' rights can be exercised regarding their personal data. AI tools may be employed to assist the recruitment process, but without replacing human judgment.