- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- il y a 2 heures
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Blend
Blend is a leading AI services provider focused on creating significant impact by integrating data science, AI, technology, and human experience. The company’s mission is to empower bold goals by combining expert human knowledge with artificial intelligence, unlocking value and fostering innovation. Blend emphasizes meaningful outcomes through people and AI collaboration.
Job Summary
The role seeks a skilled Azure Databricks Engineer adept in Python, SQL, and Apache Spark to architect, construct, and enhance scalable data processing and analytical pipelines on Microsoft Azure cloud platform. The candidate should have expertise managing large datasets, distributed data processing, and modern data engineering methodologies.
Key Responsibilities
- Architect, develop, and sustain scalable data pipelines on Azure Databricks.
- Design and implement ETL and ELT workflows using PySpark, Spark SQL, and Python.
- Optimize Spark jobs focusing on performance efficiency, cost reduction, and scalability.
- Handle structured and semi-structured data formats such as Parquet, Delta Lake, JSON, and CSV.
- Create and manage Delta Lake tables supporting ACID transactions, time travel, and schema evolution.
- Integrate Azure Databricks systems with Azure Data Lake Storage Gen2.
- Build complex SQL queries and data transformations.
- Collaborate with data scientists, analysts, and other stakeholders to enable analytics and machine learning initiatives.
- Maintain data quality through validation and monitoring procedures.
- Adhere to best practices regarding security, access control, and governance within the Azure environment.
Qualifications and Skills
- A minimum of 4 years of experience in Data Engineering.
- Proficiency and hands-on experience with Azure Databricks platform.
- Strong programming skills in Python specifically for data processing.
- Advanced knowledge of SQL including joins, window functions, and performance tuning.
- Practical experience with Apache Spark and PySpark frameworks.
- Expertise with Delta Lake storage and its features.
- Familiarity with Azure Data Lake Storage Gen2.
- Understanding of distributed computing concepts.
- Experience utilizing Git for version control.
Additional Skills and Knowledge
- Experience working with Azure Data Factory for data orchestration.
- Exposure to continuous integration and continuous deployment pipelines using tools such as Azure DevOps and GitHub Actions.
- Basic knowledge of data modeling principles.
- Awareness of cloud security practices, especially Role-Based Access Control (RBAC) within Azure.
- Experience with streaming data technologies including Spark Structured Streaming, Event Hub, and Kafka.