Senior Platform Engineer - Network Infrastructure
Hyderabad, Telangana, India · ಪೂರ್ಣ ಸಮಯ
ಅರ್ಜಿ ಸಲ್ಲಿಸುವವರಲ್ಲಿ ಮೊದಲಿಗರಾಗಿರಿ
- ಅನುಭವ
- 8+ ವರ್ಷಗಳು
- ಸಂಬಳ
- —
- ತೆರೆಯುವಿಕೆಗಳು
- 1
- ಪೋಸ್ಟ್ ಮಾಡಲಾಗಿದೆ
- 23 ಗಂಟೆಗಳು ಹಿಂದೆ
- ಕೆಲಸದ ಮೋಡ್
- ಕಚೇರಿಯಲ್ಲಿ
- ವಿದ್ಯಾಭ್ಯಾಸ
- ಯಾವುದೇ ಪದವೀಧರರು
- ಅರ್ಹತೆ
- Candidates with any graduate degree are eligible to apply.
- ಪುನರಾರಂಭ
- ಅರ್ಜಿ ಸಲ್ಲಿಸಲು ಕಡ್ಡಾಯ
ನೀವು ಎಲ್ಲಿ ಕೆಲಸ ಮಾಡುತ್ತೀರಿ
ಕೆಲಸದ ವಿವರ
About the Role
Join NVIDIA's Cloud Foundations Reliability (CFR) team, part of the Global Network Infrastructure (GNI) group, focused on deploying, integrating, and operating a Kubernetes-based platform that underpins NVIDIA's global network across data centers, colocation centers, and cloud settings. The team manages the platform's architecture and lifecycle, including cluster provisioning, upgrades, GitOps delivery, observability, and capacity management. The goal is to develop software and automation to standardize deployment, scalability, and management of network platforms and services across diverse environments.
Key Responsibilities
- Design, develop, and maintain the Kubernetes platform enabling network automation, telemetry, and operations spanning data centers, colocation, and cloud environments.
- Manage the lifecycle of GNI Kubernetes clusters including onboarding, upgrades, capacity planning, availability, and disaster recovery procedures.
- Create production-grade software and automation tools for cluster provisioning, validation, upgrades, remediation, and multi-cluster delivery using GitOps methodologies.
- Provide production-level support for network services hosted on the platform, collaborating with respective service owners on architecture and features.
- Troubleshoot complex platform and service failures involving control-plane health, networking, storage, workload scheduling, placement, and multi-cluster dependencies, driving issues to resolution.
- Establish standards for production readiness and observability for both platform and hosted services, including health monitoring, capacity metrics, alerting, runbooks, and recovery plans.
- Participate in on-call rotations, including evenings and weekends, leading incident response and follow-up to ensure corrective measures are executed.
Qualifications
- Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- More than 8 years of experience developing or operating production Kubernetes platforms, network infrastructure, or distributed systems.
- Extensive expertise in Kubernetes operations at scale, including lifecycle management, upgrades, networking, storage, and recovery processes.
- Proficiency with at least one programming language such as Go or Python.
- Hands-on experience with GitOps practices, infrastructure as code, CI/CD pipelines, and automated production deployments.
- Experience deploying and maintaining network automation and telemetry applications on Kubernetes platforms.
- Familiarity with handling on-call duties, incident response, root cause analysis, and implementing corrective actions.
Preferred Skills and Experience
- Strong understanding of IP routing, data center fabrics, and cloud networking architectures.
- Experience managing large, multi-region Kubernetes fleets, including performing fleet-wide upgrades and recovery strategies.
- Practical experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle management, machine remediation, and upgrades.
- Development of Kubernetes controllers or operators in Go, utilizing custom resources and reconciliation patterns.
- Track record of designing or supporting network automation and telemetry services at global scale on Kubernetes.
- Contributions to open-source Kubernetes infrastructure projects such as Cluster API or Metal3.
Additional Information
NVIDIA's deep learning platforms significantly influence various industries and institutions worldwide. We seek passionate, dedicated, and innovative professionals ready to tackle challenging opportunities in deep learning cloud solutions. NVIDIA is known as a leading technology employer with some of the most advanced and motivated talent in the field. If you thrive on creativity, autonomy, and challenges, this could be an ideal fit.
Eligibility: Applications are welcome from any graduates meeting the qualifications. Please verify all job details and current availability with NVIDIA directly before applying.