Staff Site Reliability Engineer
Previously generated summary being refreshed
Replit is hiring a Staff Site Reliability Engineer to work remotely across several European countries. The role focuses on designing observability, defining SLOs/SLIs, leading incident response, and automating infrastructure using tools such as Terraform, Kubernetes, and GCP. Candidates must have 8‑10 years of SRE or related experience, strong Python or Go programming skills, and deep knowledge of distributed systems and container orchestration. The position includes mentoring responsibilities and contributes to the reliability culture across the engineering organization.
Key details
- Remote work available in the UK, Italy, Netherlands, Ireland, France, and other European locations
- Staff‑level SRE position requiring 8‑10 years of experience
- Must program in Python or Go and work with Terraform, Kubernetes, and GCP
- Responsible for observability, SLO/SLI definition, incident leadership, and automation
- Includes mentorship and guidance for engineering teams
