We are seeking a Senior DevOps Engineer to serve as an escalation point for advanced incident management, lead root cause analysis efforts, and drive the engineering and optimization of CI/CD pipelines and infrastructure-as-code across complex, multi-platform environments.
Responsibilities
- Serve as the escalation point for advanced incident management, performing deep-dive troubleshooting of infrastructure, containers, and cloud services to resolve complex or prolonged incidents
- Lead Root Cause Analysis for major or recurring issues, document findings, and implement permanent corrective actions to reduce recurrence
- Troubleshoot pipeline failures, maintain infrastructure-as-code using Terraform and Helm, and develop and optimize the CI/CD toolchain
- Plan and execute non-standard, high-risk, or architectural changes, including cluster migrations and major upgrades
- Analyze system performance, adjust resources, and tweak configurations to ensure stability and cost optimization
- Create and update Standard Operating Procedures and Runbooks, and conduct training sessions
- Manage and maintain Windows and Linux environments across multiple business-critical applications
- Configure and support IIS, HAProxy, RabbitMQ, and Active Directory infrastructure
- Handle Security TLS certificates management across various systems and platforms
- Support and maintain Kubernetes clusters and AWS cloud infrastructure
Requirements
- 5+ years of experience in DevOps, SRE, or infrastructure engineering roles
- Expertise in Windows and Linux administration, PowerShell, and Bash scripting
- Proficiency in Kubernetes, Terraform, and Helm for container orchestration and infrastructure-as-code
- Knowledge of AWS services including IAM, EKS, ASG, ALB/NLB, and Route53
- Background in Ansible for configuration management and automation
- Familiarity with IIS, HAProxy, and RabbitMQ administration
- Understanding of Active Directory and Security TLS certificate management
- Experience with PowerBI, DWH, and MS SQL for data and reporting infrastructure
- Skills in CI/CD pipeline development, troubleshooting, and optimization
- Competency in performance tuning, root cause analysis, and incident management practices
- Proficiency in English at a B2+ level
We offer
- CONTINUOUS UPSKILLING, LEARNING \& DEVELOPMENT
- Diversity of tasks and projects
- Assessment center for objective review of competency level
- Personal development plan
- Mentoring programs and leadership development
- Certification and professional development support
- Access to learning platforms including more than 2,500 internal courses
- English courses taught by certified teachers
- CORPORATE BENEFITS
- Extra leave days
- Referral bonuses
- COMPENSATION PACKAGE
- Competitive compensation paid in USD
- Regular salary and performance reviews
- MEDICAL \& HEALTHCARE
- Private health insurance
- Well-being events
- WORKING ENVIRONMENT
- Recreation areas and kitchens
- Tea, coffee and snacks
- Sports equipment and game consoles
- IT Equipment
- Microsoft’s Software Assurance Home Use Program (HUP)
Please note that our Talent Attraction Team reviews applications and CVs submitted in English.
EPAM is a global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. We deliver globally and engage locally, making the future real for clients, partners, and employees. We are proud to be recognised by Forbes, Eurostaffs, Newsweek, Time Magazine, Great Place to Work and kununu as a Most Loved Workplace around the world.