AWS Data Engineer (PySpark)
Job Description
Good to have: AWS, Pyspark, Terraform + Public Sector Key Responsibilities • Design, develop, and maintain scalable data pipelines using PySpark on AWS. • Build and orchestrate ETL/ELT workflows using AWS Glue and AWS Step Functions. • Develop serverless applications and automation using AWS Lambda. • Write clean, efficient, and maintainable Python/PySpark code following engineering best practices. • Provision and manage cloud infrastructure using Terraform (Infrastructure as Code). • Implement and maintain CI/CD pipelines to automate code deployment, testing, and infrastructure changes. • Monitor, troubleshoot, and optimize data pipelines for performance, reliability, and cost efficiency. • Collaborate with business stakeholders to deliver data solutions. • Follow DevOps, security, and coding standards throughout the engagement.
Required Skills • Strong hands-on experience with PySpark and Python for data engineering. • Experience developing ETL pipelines using AWS Glue. • Proficiency with AWS Step Functions for workflow orchestration. • Experience building serverless solutions using AWS Lambda. • Hands-on experience with Terraform for Infrastructure as Code (IaC). • Experience implementing CI/CD pipelines using tools such as GitLab, GitHub Actions, Jenkins, or similar. • Good understanding of AWS services, data lakes, IAM, S3, CloudWatch, and monitoring.