MLOps Engineer
Abu Dhabi Emirate, United Arab Emirates · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- há 7 horas
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Inception42
Inception42, a G42 company based in Abu Dhabi, UAE, leads innovation in AI by delivering both domain-specific and versatile AI products. The company acts as the intelligence core within G42, converting data and computational infrastructure into tangible AI solutions with a commitment to societal benefits.
Role Overview
We seek an experienced MLOps Engineer to develop and maintain robust pipelines and systems that facilitate the transition of machine learning models from development stages to stable production environments. This position involves collaborating with data scientists, machine learning and software engineers, architects, and product teams to ensure operational excellence and scalability.
Key Responsibilities
- Design and manage comprehensive machine learning pipelines including stages like data prep, training, validation, packaging, deployment, and ongoing production monitoring.
- Develop and maintain CI/CD workflows tailored for ML systems featuring automated testing, environment promotion, rollback capabilities, and reproducibility.
- Oversee model lifecycle management, encompassing experiment tracking, version control, deployment governance, and compliance checks.
- Deploy and monitor model-serving and inference operations within containerized, cloud-native ecosystems across various environments.
- Create reusable components and standards to streamline the migration of models from research phases into production.
- Implement instrumentation for monitoring model quality, detecting data and concept drift, latency, throughput, availability, and infrastructure health metrics.
- Automate infrastructure provisioning leveraging infrastructure-as-code practices and manage access controls and security requirements effectively.
- Diagnose and optimize pipelines, deployment paths, and platform services to enhance reliability, scalability, efficiency, and cost-effectiveness.
- Integrate security, compliance, and data governance into platforms and deployment workflows.
- Collaborate with cross-functional teams to resolve production-level issues, enhance documentation, and evolve MLOps capabilities.
Candidate Profile
- Strong foundation in software engineering, machine learning systems, cloud infrastructure, automation, and operational management.
- Proven experience in building and maintaining production ML or data platforms with a focus on deployment quality, security, and reliability.
- Proficiency in developing CI/CD pipelines using Python, shell scripting, or similar technologies.
- Hands-on knowledge with containerization and orchestration tools like Docker and Kubernetes, including troubleshooting across application and infrastructure layers.
- Familiarity with Azure cloud services (e.g., Azure Machine Learning, AKS, Azure DevOps) or equivalent cloud platforms for ML and application workloads.
- Experience with ML lifecycle management tools such as MLflow, Kubeflow, Airflow, DVC, or analogous frameworks.
- Strong skills in monitoring and debugging model behavior, data quality, inference efficiency, distributed services, and infrastructure health.
- Excellent communication, pragmatic decision-making, and the ability to collaborate across diverse technical and product teams in ambiguous situations.
Additional Desirable Skills
- Experience with generative AI or large language model deployments, including prompt/version management and high-performance GPU inference environments.
- Understanding of distributed data platforms, search systems, various database technologies (SQL/NoSQL), time-series storage, feature stores, and large-scale training infrastructure.
- Familiarity with infrastructure automation tools such as Terraform or Bicep and observability tools including Prometheus, Grafana, ELK stack, Datadog, or Langfuse.