EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are looking for a Senior Data Engineer to design and deliver scalable data pipelines and high-performance analytical solutions using SQL/BigQuery, Spark/PySpark, and Python on Google Cloud. This role focuses on building reliable, cloud-native data products that enable advanced reporting, analytics, and decision-making across the organization.
Responsibilities
- Design and deliver batch and real-time ETL/ELT pipelines across cloud environments to support analytics and reporting
- Write and optimize advanced SQL transformations and build performant, cost-efficient BigQuery data models
- Implement scalable data processing solutions using Python and PySpark, ensuring maintainable and high-quality code
- Build robust data models and apply validation practices to maintain accuracy and reliability
- Use GCP services such as BigQuery, Dataflow, Cloud Composer, Pub/Sub, and GCS to build and operate modern data platforms
- Troubleshoot complex pipeline issues and continuously improve compute, storage, and processing performance
- Collaborate with engineering, analytics, and business teams while contributing to CI/CD, code reviews, and testing standards
Requirements
- 5+ years of experience as a Data Engineer, building scalable data pipelines and working with cloud-based data ecosystems
- Expertise in SQL and hands-on experience building performant datasets in BigQuery or similar cloud data warehouses
- Proficiency in Python and PySpark for scalable data processing in distributed environments
- Understanding of data modeling, ELT/ETL patterns, and data quality best practices
- Familiarity with Google Cloud Platform, particularly BigQuery, Dataflow, and Cloud Composer, GCS, or equivalent cloud data services
- Background in building scalable data pipelines, both batch and near real-time, in a cloud-native environment
- Proficiency with version control, CI/CD pipelines, and automated testing frameworks
- Capability to troubleshoot and optimize performance across compute, storage, and processing layers
Nice to have
- Experience with Infrastructure-as-Code, including Terraform, Ansible, and Chef
- Knowledge of shell scripting
- Experience in financial services or regulated environments
We offer
- Full access to cutting-edge tools and technologies
- Competitive compensation depending on experience and skills
- All-around Social package: professional \& soft skills training, medical \& family care programs, sports
- Free English classes
- Unlimited access to Eurostaffs learning solutions
- Continuous experience exchange with experts and professionals worldwide
- Friendly team and comfortable working environment
- Engineering, corporate, and social events within and outside the Company
- Flexible working schedule
- Opportunities for self-realization