Job description
The Principal Data Engineer is responsible for designing, building, and leading enterprise-scale data platforms that enable analytics, business intelligence, AI/ML, and operational reporting. This role provides technical leadership for data engineering initiatives, defines data architecture standards, mentors engineering teams, and ensures scalable, secure, and high-performance data solutions across cloud and on-premises environments. The ideal candidate has extensive experience in modern data platforms, cloud technologies, data warehousing, ETL/ELT frameworks, and distributed data processing.
Responsibilities
- Design and implement scalable, high-performance data platforms and pipelines for enterprise data integration.
- Lead the architecture and development of cloud-based data solutions using modern data engineering technologies.
- Define data engineering standards, best practices, and governance frameworks across the organization.
- Build and optimize ETL/ELT processes for structured and unstructured data from multiple source systems.
- Develop and maintain data warehouses, data lakes, and lakehouse architectures.
- Optimize data models, SQL queries, and processing frameworks for performance and scalability.
- Collaborate with business stakeholders, architects, analytics teams, and data scientists to deliver reliable data solutions.
- Ensure data quality, security, compliance, and governance throughout the data lifecycle.
- Lead technical design reviews, code reviews, and solution architecture discussions.
- Mentor and guide senior and junior data engineers while driving engineering excellence.
- Evaluate emerging technologies and recommend improvements to the organization's data ecosystem.
- Support CI/CD implementation, DevOps practices, monitoring, and automation for data platforms.
- Troubleshoot complex production issues and drive continuous platform optimization.
Requirements
- 10–15+ years of experience in Data Engineering, Data Warehousing, or Big Data technologies.
- Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Strong expertise in SQL and programming languages such as Python, Scala, or Java.
- Extensive experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform.
- Hands-on experience with modern data platforms including Databricks, Snowflake, Microsoft Fabric, Synapse Analytics, or Redshift.
- Strong knowledge of distributed processing frameworks such as Apache Spark and PySpark.
- Experience building scalable ETL/ELT pipelines using Azure Data Factory, Apache Airflow, Informatica, or similar tools.
- Expertise in data modeling, dimensional modeling, and enterprise data warehouse design.
- Experience with Delta Lake, Lakehouse architecture, and data governance frameworks.
- Knowledge of streaming technologies such as Kafka or Event Hubs is preferred.
- Experience with DevOps, Git, CI/CD pipelines, Docker, and Kubernetes is an advantage.
- Strong understanding of data security, privacy, and compliance best practices.
- Excellent analytical, communication, stakeholder management, and leadership skills.
This job post has been translated by AI and may contain minor differences or errors.