LinkedIn y terceros utilizan cookies imprescindibles y opcionales para ofrecer, proteger, analizar y mejorar nuestros servicios, y para mostrarte publicidad relevante (incluidos anuncios profesionales y de empleo) dentro y fuera de LinkedIn. Consulta más información en nuestra Política de cookies.
Selecciona Aceptar para consentir o Rechazar para denegar las cookies no imprescindibles para este uso. Puedes actualizar tus preferencias en cualquier momento en tus ajustes.
We are looking for a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team.
In this role, you will bridge the gap between software development and systems operations. You will apply software engineering principles to automate our operations, scale our infrastructure, and ensure our systems are highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and helping us deploy software rapidly and safely.
Responsibilities
Design, build, and maintain cloud infrastructure using modern Infrastructure as Code (IaC) practices such as Terraform or CloudFormation
Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, or Datadog
Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure system reliability
Respond to production incidents and lead troubleshooting efforts to restore services quickly
Conduct blameless post-mortems to identify root causes and prevent recurring issues
Partner with software developers to optimize system performance and plan for capacity needs
Ensure services can scale effectively to handle growth and traffic spikes
Requirements
A minimum of 3 years of relevant experience
Strong background in systems administration, DevOps, or systems-focused software development
Proficiency in at least one scripting or programming language, such as Python, Bash, Go, or Rust
Solid experience with public cloud providers, including AWS, Azure, or GCP
Hands-on experience with containerization tools such as Docker and Kubernetes
Deep understanding of Linux/Unix administration and core networking fundamentals, including TCP/IP, DNS, HTTP, and SSL/TLS
A genuine passion for automation, reducing manual toil, and building resilient systems that fail gracefully
Experience within Financial Services, Insurance, or Retail industries
Strong written and spoken proficiency in English at a C1 level or higher
We offer
International projects with top brands
Work with global teams of highly skilled, diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling, reskilling and certification courses
Unlimited access to the LinkedIn Learning library and 22,000+ courses
Global career opportunities
Volunteer and community involvement opportunities
EPAM Employee Groups
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Nivel de antigüedad
Intermedio
Tipo de empleo
Jornada completa
Función laboral
Ingeniería, Tecnología de la información y Desarrollo empresarial
Sectores
Desarrollo de software, Servicios y consultoría de TI y Organización de viajes
Las recomendaciones duplican tus probabilidades de conseguir una entrevista con EPAM Systems