You'll own the infrastructure that tells Luma whether its models are getting better. As a Research Engineer on Evaluations, you'll build the pipelines, metrics, and automated systems that close the loop between model output, measurement, and improvement, so every model decision is grounded in rigorous, consistent signal. This sits across research, engineering, and product, and it's as much a systems job as an ML one. You'll be turning human evaluation frameworks into automated ones and wiring the results straight back into training. It fits someone who's happy owning both the ML and the infrastructure around it. If you only want to work on models and not the pipelines that measure them, this isn't the seat. What You'll Own Design and build scalable pipelines for automated evaluation of generative models across image, video, text, and audio. Develop metrics and evaluation models that capture fidelity, coherence, temporal consistency, and alignment with human intent. Integrate evaluation signals into training loops, including reinforcement learning and reward modeling, to keep improving the models. Build the infrastructure for large-scale regression testing, benchmarking, and monitoring of multimodal models. Work with researchers running human studies to translate their frameworks into automated or semi-automated systems. Maintain the dashboards, reporting, and alerting that surface evaluation results to the teams that need them. First 90 Days One way the first 90 could unfold. Days 1–30 — Immerse & Diagnose: Map how models are evaluated today, where the gaps and manual steps are, and which signal matters most to model development. Days 30–60 — Ship & Validate: Stand up an automated evaluation pipeline for one modality and prove it catches real regressions the current setup misses. Days 60–90 — Scale & Systemize: Extend it into training loops and monitoring so evaluation runs continuously, not as a one-off. What You Bring 5+ years building ML evaluation systems, model pipelines, or large-scale infrastructure. Master's or PhD in Computer Science, Machine Learning, or a related field, or equivalent industry experience. Hands-on experience with visual data (image and/or video) in evaluation, modeling, or data prep. Proficiency in Python and an ML framework (PyTorch, JAX, or TensorFlow). Strong ML background with generative models (diffusion, LLMs, multimodal architectures). Strong software engineering: CI/CD, testing, data pipelines, distributed systems. Familiarity with human-in-the-loop evaluation and how to scale it with automation. Nice to Have Experience with reinforcement learning or reward modeling. Prior work on perceptual metrics, multimodal benchmarks, or retrieval-based evaluation. Background in large-scale model training or evaluation infrastructure. Familiarity with creative media workflows (film, VFX, animation, digital art). Contributions to open-source evaluation libraries or benchmarks. About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
Every tech & IT company hiring across India — with AI match scores — on one live map.
Open the map →