Big Data Engineer – Java, Apache Spark, AWS EMR, Lambda, EKS & Airflow
Description
Job Summary
Synechron is seeking a Big Data Engineer with 5+ years of experience in designing, developing and deploying scalable and reliable data-processing solutions on AWS using Java and Apache Spark. The role will focus on implementing big data processing pipelines using Apache Spark on AWS EMR, developing and deploying serverless applications using AWS Lambda, utilizing Amazon EKS for container orchestration and microservices management, and designing workflow orchestration using Apache Airflow. The position will contribute to business objectives by delivering reliable, scalable and maintainable data-processing solutions that support efficient data operations and application services.
Software Requirements
Required
- Java: Strong hands-on experience in Java development.
- Apache Spark: Proficiency in Apache Spark for distributed data processing.
- AWS EMR: Experience implementing big data processing pipelines using Apache Spark on AWS EMR.
- AWS Lambda: Experience developing and deploying serverless applications using AWS Lambda.
- AWS Service Integration: Experience integrating AWS Lambda with other AWS services.
- Amazon EKS: Experience utilizing Amazon EKS for container orchestration and microservices management.
- Apache Airflow: Experience designing and implementing workflow orchestration for complex data pipelines.
- AWS: Experience designing, developing and deploying scalable and reliable data-processing solutions on AWS using Java and Spark.
- Project-Supported Versions: Ability to work with the versions of Java, Apache Spark, AWS services and related tools supported by the project.
Preferred
- No additional preferred software skills are specified in the requirements. Additional experience relevant to Java, Apache Spark, AWS, AWS EMR, AWS Lambda, Amazon EKS or Apache Airflow may be considered beneficial.
Overall Responsibilities
- Design, develop and deploy scalable and reliable data-processing solutions on AWS using Java and Apache Spark.
- Implement and maintain big data processing pipelines using Apache Spark on AWS EMR.
- Develop and deploy serverless applications using AWS Lambda and integrate them with other AWS services.
- Utilize Amazon EKS for container orchestration and microservices management.
- Design, implement and maintain Apache Airflow workflows for complex data pipelines.
- Develop maintainable Java applications and data-processing components aligned with technical requirements.
- Contribute to solution design, technical discussions and implementation decisions within agreed project standards.
- Test, troubleshoot and resolve application, pipeline, integration and deployment issues.
- Support reliable operation of data-processing solutions across relevant development, testing and production environments.
- Monitor application and pipeline behavior and contribute to improvements in scalability, reliability and maintainability.
- Collaborate with engineering, architecture, QA and delivery teams to clarify requirements and coordinate implementation activities.
- Contribute to code reviews, technical documentation, release activities and knowledge sharing.
- Apply appropriate data protection and cloud security practices during development and deployment.
- Use AWS resources efficiently and consider sustainable practices that reduce unnecessary compute, storage and network consumption.
- Deliver assigned work in accordance with agreed priorities, timelines and quality expectations.
Technical Skills (By Category)
Programming Languages
Essential
- Strong hands-on experience in Java development.
- Ability to develop clean, modular, reusable and maintainable Java code.
- Experience using Java to develop data-processing solutions and supporting application components.
- Ability to debug Java code and resolve functional, integration and deployment issues.
Preferred
- Additional experience with languages used in data engineering may be considered beneficial.
Databases and Data Management
Essential
- Experience implementing big data processing pipelines.
- Understanding of distributed data processing using Apache Spark.
- Ability to design data-processing workflows that support reliability and scalability.
- Ability to manage pipeline execution, dependencies and processing failures.
Preferred
- Additional experience with data storage, transformation, validation or integration may be considered beneficial.
Cloud Technologies
Essential
- Experience designing, developing and deploying data-processing solutions on AWS.
- Experience with AWS EMR for Apache Spark-based big data processing.
- Experience with AWS Lambda for serverless application development and deployment.
- Experience integrating AWS Lambda with other AWS services.
- Experience with Amazon EKS for container orchestration and microservices management.
- Experience using project-supported AWS service configurations and versions.
Preferred
- Additional experience relevant to AWS-based data-processing and application deployment may be considered beneficial.
Frameworks and Libraries
Essential
- Apache Spark: Proficiency in distributed data processing.
- Apache Airflow: Experience designing and implementing workflow orchestration for complex data pipelines.
- Experience integrating Apache Spark with AWS EMR.
- Experience using frameworks and libraries compatible with project-supported Java, Spark and AWS versions.
Preferred
- Additional experience developing reusable data-processing components may be considered beneficial.
Development Tools and Methodologies
Essential
- Experience with the design, development and deployment lifecycle for data-processing solutions.
- Experience deploying and maintaining applications and pipelines in AWS environments.
- Ability to troubleshoot data pipelines, serverless applications, containerized services and workflow processes.
- Ability to document technical solutions, workflows and deployment-related activities.
- Ability to collaborate with relevant technical and delivery stakeholders.
Preferred
- No additional development tools or methodologies are specified in the requirements.
- Experience with automated testing, deployment or monitoring practices may be considered beneficial.
Security Protocols
Essential
- Understanding of secure development practices for cloud-based applications and data-processing solutions.
- Awareness of appropriate access control for AWS services and deployed applications.
- Ability to protect data during processing, transmission and storage.
- Ability to manage application configuration and credentials securely.
- Ability to apply secure logging and error-handling practices.
Preferred
- No specific additional security protocols are specified in the requirements.
- Experience with cloud security controls for serverless, containerized and distributed data-processing solutions may be considered beneficial.
Experience Requirements
- 5+ years of experience in big data engineering, Java development, cloud engineering or a related technical role.
- Experience designing, developing and deploying scalable and reliable data-processing solutions on AWS using Java and Apache Spark.
- Experience implementing big data processing pipelines using Apache Spark on AWS EMR.
- Experience developing and deploying serverless applications using AWS Lambda and integrating them with other AWS services.
- Experience utilizing Amazon EKS for container orchestration and microservices management.
- Experience designing and implementing workflow orchestration using Apache Airflow for complex data pipelines.
- Strong hands-on experience in Java development and proficiency in Apache Spark for distributed data processing.
- Experience with AWS services including EMR, Lambda, EKS and Airflow.
- Experience supporting the deployment, troubleshooting and maintenance of AWS-based dat