- Design and implement scalable data pipelines using Python, PySpark, and boto3
- Integrate with Amazon Redshift for data warehousing and perform data ingestion and transformation tasks
- Work with Amazon S3 to architect and manage data lakes
- Build and manage batch processing jobs using AWS Glue, EMR, and Lambda
- Support streaming data workflows using Amazon Kinesis
- Leverage AWS Secrets Manager and KMS for secure data handling and encryption
- Write effective SQL for data querying, transformations, and performance tuning
- Collaborate with data engineering, product, and DevOps teams to deliver end-to-end solutions
- Participate in code reviews, debugging, and testing activities to ensure high-quality software delivery
- Use Git for version control and follow best practices in CI/CD
- Continuously explore and recommend improvements to enhance system performance, scalability,
and maintainability
Required Skills:
- Advanced proficiency in Python, including libraries such as Pandas, PySpark, and boto3
- Experience designing workflows in Apache Airflow (MWAA)
- Hands-on experience with Amazon Redshift for analytics and data warehousing
- Solid SQL skills and experience with data modeling and transformation
- Comfortable with Amazon S3, including setting up data lake architecture
- Familiar with AWS core services: EC2, S3, IAM, Lambda
- Experience using AWS Secrets Manager and AWS KMS for secure data access