About the Role
We are looking for a hands-on Lead Data Engineer to drive the technical evolution of our data platform. This role combines approximately 50% hands-on engineering with 50% architecture, technical leadership, and mentorship.
You will be responsible for shaping a high-throughput data infrastructure, building production-grade pipelines and data assets, improving platform scalability and reliability, and helping a growing engineering team raise its technical standards.
Key Responsibilities
1. Data Platform Architecture & Scalability
- Define and drive the technical architecture of the company‘s streaming and batch data infrastructure as data volumes and platform requirements continue to scale.
- Identify and resolve performance bottlenecks across data storage, processing, and querying layers.
- Optimize data models, query execution, and storage strategies to achieve the right balance between performance, scalability, and infrastructure costs.
- Establish architectural standards and technical direction for the broader data platform.
2. Data Pipeline & Identity Engineering
- Design and implement reliable, high-throughput ETL pipelines using the Medallion architecture to transform raw data into trusted, analytics-ready datasets.
- Build fault-tolerant data processing workflows capable of handling large-scale data volumes.
- Design and evolve identity graph capabilities that consolidate fragmented consumer attributes, visitor information, leads, and customer identities into a unified identity layer.
- Ensure core data products remain reliable, scalable, and fit for downstream analytics and product use cases.
3. ML Data Engineering & MLOps
- Develop and maintain feature pipelines and feature stores that support production AI/ML models, including use cases such as identity resolution and propensity scoring.
- Work closely with Data Science teams to productionize models and integrate them into scalable data workflows.
- Support model deployment, infrastructure scaling, monitoring, and retraining processes.
- Ensure data pipelines provide reliable and production-ready inputs for ML systems.
4. Data Governance & Engineering Excellence
- Establish data governance standards covering data ownership, lineage, cataloging, access control, and compliance.
- Implement frameworks and best practices for data quality, pipeline testing, CI/CD, and observability.
- Improve the reliability and maintainability of the company‘s data infrastructure through automation and engineering best practices.
- Mentor junior and mid-level engineers and contribute to a strong culture of technical ownership and continuous improvement.
Requirements
- 8+ years of professional experience in Data Engineering, with demonstrated experience leading technical initiatives or mentoring engineers.
- Strong programming skills in Python; experience with Node.js is a plus.
- Advanced SQL skills and strong expertise in data modeling for large-scale analytical environments.
- Extensive production experience with cloud-based data platforms, including AWS (S3, Kinesis, Athena, Redshift, Lambda) or equivalent GCP services such as BigQuery, Dataflow, Pub/Sub, and Cloud Storage.
- Hands-on experience with ClickHouse or another large-scale columnar/OLAP database.
- Strong experience designing and operating data workflows with Airflow or similar orchestration frameworks.
- Solid understanding of Lakehouse architecture and hands-on experience implementing Medallion data layering.
- Experience developing data pipelines that support ML/AI use cases, including feature pipelines or feature stores.
- Strong English communication skills and the ability to collaborate effectively with distributed, multi-country engineering teams.
Nice to Have
- Background in MarTech/AdTech, particularly identity resolution, first-party cookie data, or digital marketing platforms.
- Exposure to Data Science / MLOps technologies and frameworks.
- Experience working in an early-stage or rapidly scaling B2B SaaS environment with a strong ownership culture.