- Experience
- 1–3 yrs
- Salary
- —
- Openings
- 1
- Posted
- منذ 11 ساعة
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Job description
About the Role
Kerry Consulting is a rapidly expanding tech firm focused on creating scalable data platforms to enhance AI capabilities, analytics, and next-generation digital products. We are currently recruiting a Junior Data Engineer to aid in building dependable data pipelines and modern data infrastructure. This role offers a valuable chance for an early-stage data professional to collaborate with seasoned data engineers, software developers, and data scientists while developing enterprise-grade data platforms that support AI-driven solutions.
Key Responsibilities
You will be accountable for designing, developing, and maintaining ETL and ELT data pipelines that aggregate, transform, and deliver data from diverse sources into enterprise data systems. Collaborating closely with data engineers, scientists, and app development teams, you will guarantee the delivery of high-quality, reliable, and scalable data flows backing reporting, analytics, and AI projects. Tasks will include writing and optimizing SQL queries, developing and tuning PySpark processing jobs, performing robust data validation and quality assurance, troubleshooting pipeline malfunctions, and contributing to enhancements in data reliability and system performance. Additional duties encompass supporting data modeling, documentation, monitoring, automation, and integrating modern cloud-based data engineering methodologies and tools.
Candidate Profile and Requirements
The ideal candidate will have between one to three years of experience in data engineering, database programming, or ETL development within a tech or data-centric organization. Proficiency in SQL and practical experience designing and managing ETL/ELT pipelines are essential. Candidates must possess hands-on knowledge of PySpark, including building and refining distributed data processing workflows for substantial datasets.
Experience using Python, relational database management systems, and contemporary data platforms such as Databricks, Snowflake, Azure Data Factory, Apache Airflow, AWS Glue, or Apache Spark is strongly preferred. Familiarity with cloud environments (AWS, Azure, Google Cloud), data modeling techniques, and workflow orchestration tools represents a significant advantage. Candidates should demonstrate excellent analytical and problem-solving capabilities, meticulous attention to data quality, and the aptitude to work collaboratively with cross-functional teams to deliver scalable, reliable data solutions.