Staff Site Reliability Engineer
Manchester, England, United Kingdom · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- GBP 124,000 – GBP 141,000 / year
- Openings
- 1
- Posted
- منذ 3 ساعات
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Job description
About Obsidian Security
Established in 2017, Obsidian Security addresses a critical need by securing SaaS applications such as Microsoft 365 and Salesforce where modern businesses operate. Supported by top investors like Greylock and Norwest Venture Partners, Obsidian has developed a comprehensive SaaS security platform aimed at minimizing risk, detecting and responding to threats, and preventing breaches at their source. The leadership team includes pioneers from industry leaders like CrowdStrike, Okta, and Cylance. Trusted by over 200 global organizations including Snowflake and T-Mobile, Obsidian secures numerous Fortune 1000 and Global 2000 companies across multiple continents. With rapid global expansion, a growing partner ecosystem including SentinelOne and Google Cloud, and plans for an IPO, Obsidian is shaping the future of SaaS security.
Role Summary
As a Staff Site Reliability Engineer (SRE), you will lead and shape the company’s reliability vision for a complex multi-tenant SaaS platform serving enterprise and financial clients. Acting as a strategic partner to DevOps and Platform Engineering teams, you will develop and execute a unified reliability strategy that supports scalability and resilience across the organization. Your primary objective is to ensure system issues are detected, diagnosed, and communicated promptly to prevent customer impact.
Key Responsibilities
- Define and implement long-term reliability strategies and architectures that span multiple services, establishing comprehensive system visibility and resilience frameworks.
- Collaborate cross-functionally to integrate reliability practices, standardize service-level indicators (SLI) and objectives (SLO), and provide expert technical escalation support.
- Develop advanced detection and observability tools, including anomaly detection and connector health monitoring, and promote self-service observability capabilities.
- Lead incident management by designing a tiered communication strategy, enhancing response protocols, and conducting postmortems to improve system reliability and customer confidence.
- Contribute technically by architecting, monitoring, and troubleshooting distributed systems and data pipelines hands-on.
Required Qualifications
- Minimum of five years in Site Reliability Engineering, Production Engineering, or similar roles.
- At least three years in senior or technical leadership roles with broad organizational impact.
- Expertise with cloud platforms such as AWS and/or Google Cloud Platform.
- Proficient in Kubernetes and Helm for orchestration and deployment.
- Experience with observability tools like Prometheus, Grafana, or equivalents.
- Familiarity with continuous integration and deployment systems such as GitLab CI/CD and ArgoCD.
- Proven ability to architect and scale reliability systems in multi-tenant SaaS environments.
- Strong debugging skills and systems thinking related to distributed microservices and legacy systems.
- Track record of leading initiatives to enhance incident management, detection, and system durability.
- Hands-on engineering experience building reliability infrastructure rather than only configuring it.
Preferred Qualifications
- Experience developing B2B SaaS solutions for enterprise or financial markets.
- Understanding of third-party SaaS connector architectures and data ingestion techniques.
- Background in building intelligent alerting or anomaly detection systems.
- Knowledge of designing customer-facing status pages and frameworks for incident communication.
Why Join Us?
- Shape and drive company-wide reliability strategies.
- Own creation and enhancement of cutting-edge detection and observability systems.
- Solve challenges involving complex distributed systems.
- Protect vital infrastructure serving financial customers.
Success Indicators
- Issues are identified and resolved prior to affecting customers.
- Reliability metrics are well-defined and continuously improved.
- Teams independently utilize scalable observability tools.
- Incident communication is clear and proactive, enhancing customer trust.
- Reliability becomes a distinctive competitive advantage.
Employee Benefits
Our benefits package supports employee well-being at work and beyond, including competitive pay with equity options and 401k for US employees, comprehensive health coverage including dental and vision, flexible paid time off, 12 weeks of parental or family leave, and resources for both personal and professional growth. Compensation specifics consider location and candidate qualifications and include eligibility for equity awards and potential incentive compensation. We are committed to equal opportunity employment and provide accommodations as needed.
Compensation
The base salary range for this position is between £124,000 and £141,000 GBP.