- Experience
- 6+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 7 ਘੰਟੇ ਪਹਿਲਾਂ
- Work mode
- In office
- Education
- Bachelor's or Master's in Computer Science or related engineering field
- Resume
- Required to apply
Where you'll work
Job description
Company Overview
Elemynt, part of Xora Innovation, is an early-stage startup focused on real-world applications of AI. The platform integrates advanced machine learning, high-performance simulation, and modern software engineering to expedite the creation, validation, and deployment of novel materials. Their work merges AI, physics, and extensive computational methods, tackling challenging problems with significant real impact.
Role Summary
This senior engineering position is central to building and maintaining Elemynt's infrastructure, encompassing cloud environments, core platform services, and observability tools ensuring system reliability. The platform is designed to operate across customer-chosen environments — cloud providers, on-premises compute, or hybrid setups — typically managed by customers themselves. The infrastructure you develop must be reproducible, measurable, and inherently secure, no matter where it is deployed.
Key Responsibilities
- Develop and maintain foundational cloud infrastructure components such as networking, identity & access management, Kubernetes clusters, and infrastructure-as-code frameworks supporting all services and workloads.
- Architect and implement core platform services and internal APIs that reliably operate wherever deployed.
- Manage the entire cloud deployment lifecycle, including pipelines that convert platform services into versioned, reproducible artifacts and enable phased releases with seamless rollback capabilities.
- Establish a comprehensive observability stack comprising metrics, logging, tracing, dashboards, and alerting mechanisms for rapid fault detection and diagnosis.
- Implement service-level objectives and monitor health indicators to measure reliability and identify regressions before impacting customers.
- Ensure platform foundations are portable and reproducible, allowing consistent deployment across varied environments.
- Create deployment-ready artifacts such as container images, Helm charts, and infrastructure modules to facilitate clean, repeatable installs and upgrades.
- Enhance platform security by managing secrets, certificates, and network boundaries that safeguard software and data in any deployment context.
Required Qualifications and Experience
- Bachelor's or Master's degree in Computer Science or a related engineering discipline.
- At least six years of professional experience delivering production software with extensive expertise in cloud infrastructure and platform engineering.
- Strong hands-on proficiency with Kubernetes and infrastructure-as-code tools like Terraform on major cloud platforms (AWS, GCP, or Azure).
- Experience designing and running core backend services and APIs consumed by other systems and engineers.
- Direct responsibility for cloud deployment pipelines, managing build and release workflows, versioned artifacts, and robust, reproducible rollouts.
- Practical expertise integrating observability tools (such as Prometheus, Grafana, OpenTelemetry) into live systems and utilizing them to resolve operational incidents.
- Solid software engineering foundation with experience coding in backend or systems languages including Go, Python, or Rust.
- Competence defining service-level objectives and building systems designed for observability, reliability, and reproducibility from inception.
- Ability to thrive in an early-stage startup environment, making strategic decisions balancing scope, speed, and quality under ambiguity.
Desirable Traits and Skills
- Background in developing platform components that span diverse deployment environments, including those managed by customers.
- Experience managing workloads across heterogeneous runtimes like cloud Kubernetes, HPC schedulers (e.g., Slurm), or bare-metal servers.
- Knowledge of GPU scheduling or multi-tenant cluster operations.
- Familiarity with GitOps methodologies and progressive delivery techniques using tools like ArgoCD or Flux.
- Exposure to packaging or serving machine learning models and supporting ML and data-intensive workloads on shared platforms.
- Understanding of scientific computing, simulation, or other large-scale technical computational workloads.
Location and Work Model
Open to candidates in Singapore or the United States, with work arrangements either on-site or hybrid depending on location.
Closing Note
Candidates not fulfilling every qualification but passionate about this field are encouraged to make contact.