T

AI/ML Operations Engineer

TalentFinderX

Dubai, United Arab Emirates · Full Time

Be the first to apply

Experience
7–10 yrs
Salary
Openings
1
Posted
hace 3 horas
Work mode
In office
Education
Bachelor's degree
Resume
Required to apply

Where you'll work

Job description

Job Overview

We are seeking an experienced AI/ML Ops Engineer to oversee and maintain production systems ensuring their reliability post-deployment. The role involves end-to-end ownership of deployment pipelines, data management, monitoring, and incident response with a strong focus on system stability and continuous improvement.

Company Background

Our organization is a technology-driven company specializing in blockchain-based AI solutions. We prioritize equal opportunity and diversity, fostering an inclusive working environment. The company is headquartered in India with a small team size of 1-10 employees.

Key Responsibilities

  • Manage and maintain AI/ML production environments including deployment systems and data flows.
  • Implement and improve continuous integration and deployment workflows to ensure smooth model updates.
  • Create monitoring strategies to detect anomalies early and respond effectively to incidents.
  • Analyze system behaviors to identify root causes of reliability issues and reduce recurring incidents.
  • Communicate blockers and progress transparently, working independently with sound judgment.

Candidate Requirements

  • 7 to 10 years of practical experience in AI/ML operations or related engineering roles.
  • Proficient with model deployment, CI/CD pipelines, cloud computing, data pipeline orchestration, monitoring and logging.
  • Experienced in containerization technologies and version control systems.
  • Bachelor's degree required; master's degree preferred.
  • Professional proficiency in English and Arabic languages.
  • Ability and willingness to relocate for onsite work.

Success Milestones

  • Within 90 days: Gain a comprehensive understanding of system architecture, deployment, and monitoring to deliver tangible improvements.
  • Within 6 months: Assume full ownership of a service or pipeline, proactively manage issues, and reduce repeated problems.
  • Within 1 year: Enhance production reliability, elevate engineering standards, and contribute to strengthening team processes.

Challenges in the Role

Identifying subtle failure signals in production systems caused by drift, latency, stale data, or incomplete information is complex. The engineer must judiciously determine direct fixes, deeper investigations, or escalate decisions to maintain service quality.

Assessment and Equal Opportunity

The hiring process integrates AI-assisted assessments covering personality, voice communication, and cognitive skills. The employer upholds equal employment opportunity without discrimination based on demographic or personal characteristics.

Work styles they’re looking for

Clear Communication Practical Problem Solving Production Reliability Independent Execution Applied Judgment
🤖
Online · instant AI help