A global professional services organisation is expanding its AI and data engineering capabilities and is looking for experienced AI Data Engineers to join its Mumbai team.
This is a hands-on engineering role focused on building the modern lakehouse and data infrastructure that powers enterprise AI applications. You will design and develop scalable data platforms on Azure Databricks and ADLS Gen2, build production-grade ETL/ELT pipelines, and create governed, AI-ready data products supporting RAG and agentic AI workflows.
Key Responsibilities
- Design, build and evolve a modern Azure lakehouse using ADLS Gen2, Azure Databricks and Delta Lake.
- Implement medallion architecture (Bronze Silver Gold) for scalable and governed data processing.
- Build robust batch and streaming ETL/ELT pipelines using Azure Data Factory, Databricks Workflows and Structured Streaming.
- Develop reusable, metadata-driven ingestion frameworks and connectors for efficient onboarding of new data sources.
- Build curated, AI-ready data products to support RAG, vector search and agentic AI applications.
- Prepare data and content for embeddings and retrieval through Azure AI Search.
- Implement data governance, lineage, classification and access controls using Microsoft Purview and Databricks Unity Catalog.
- Optimise data platform performance and cost through partitioning, clustering, compaction, incremental processing and CDC.
- Establish data quality testing, CI/CD and pipeline observability across production workloads.
- Work closely with AI Developers, Product Owners, Cloud Architects and Governance & Risk teams on shared AI and data capabilities.
- Support secure handling and segregation of sensitive and regulated data.