
What makes this role different
Most data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.
The system you'll work on: Destiny is Pattern's automated Ads optimizer. Once a day it discovers the keywords worth buying for every eligible product, assembles a wide feature store from performance, bid-history and search-results data, runs 5 machine-learning models, and picks the bid level that hits each product group's return on ad spend (ROAS) and budget target. A second pipeline then pushes those campaign, keyword and budget edits to the marketplace Ads API every 15 minutes. It is a large, opinionated data system: a roughly 17,000-line orchestrated SQL codebase, a feature and label store several hundred columns wide, 5 model training and batch-scoring jobs, and blocking data-quality gates in front of every outward write. You would be one of the engineers who owns it end to end.
Roles and Responsibilities: Develop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse. Own and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and safe reruns. Write and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget. Extend the feature store - add new features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable. Orchestrate model training and batch inference on SageMaker from Airflow: build training and scoring datasets, manage S3 and Parquet round-trips, containerized training images, instance sizing, and loading predictions and metrics back into the warehouse. Develop and implement data auditing strategies and processes to ensure data quality - including blocking data-quality checks in front of outward writes - and set thresholds that catch bad data without needlessly halting live bidding. Identify and resolve problems in large-scale data processing workflows; maintain pipeline processes and troubleshoot failures, including on-call triage when a run breaks before market open. Guard the safety properties of an outward-writing system: idempotency, new-data detection, action validation and invalidation, and audit trails for every change pushed to marketplace. Collaborate with data scientists, advertising strategists, and platform teams to specify data requirements and provide access to data. Translate business and analytics requirements - ROAS targets, budget pacing, playbook rules, branded versus non-branded strategy - into a comprehensive data model and pipelines. Foster data expertise and own data quality for assigned areas of ownership; work with data infrastructure to triage issues and drive to resolution. Mentor and provide technical direction to other data engineers, and review their SQL and DAG changes.
What "basics of machine learning" means here: You are not expected to invent model architectures - data scientists own the modeling. You are expected to be a competent, unsupervised partner to them, which means being able to: Build training and evaluation datasets correctly - train/test splits over time, holdout windows, and a working instinct for target leakage in rolling-window features. Reason about class imbalance and resampling (many keyword-hours have no clicks), and about clamping or bounding predictions before they drive a bid. Read regression metrics - MAE, RMSE, MAPE, WMAPE - plus feature importances, and tell
Every tech & IT company hiring across India — with AI match scores — on one live map.
Open the map →