You'll own how Luma judges whether its models are actually good, past the point where numbers stop telling the story. As our Qualitative Evaluation Engineer, you'll build the frameworks that pin down believability, identity retention, and scene coherence, and turn them into insight that steers model development. This isn't a checkbox-metrics role. You're building evaluative systems that match the messiness of human perception and creative intent, and much of that framework doesn't exist yet. It fits someone who can take a fuzzy quality and define it in clear, testable terms, working shoulder to shoulder with researchers and technical artists. If you want work scored purely on quantitative dashboards, this isn't it. What You'll Own Evaluate generative model performance across diverse tasks, prompts, and modalities, and surface the failure modes, regressions, and edge cases that hurt product quality. Build and maintain qualitative evaluation frameworks that are scalable and reusable. Translate high-level product goals into concrete evaluative criteria. Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluations. Turn nuanced judgments into clear feedback that informs fine-tuning, dataset curation, and product UX. Work closely with technical artists and engineers to keep evaluations aligned with model capabilities and real use cases. First 90 Days One way the first 90 could unfold. Days 1–30 — Immerse & Diagnose: Learn the models and the creative use cases they serve, and audit how quality is judged today and where it misses. Days 30–60 — Ship & Validate: Stand up a qualitative framework for one high-priority capability and run it on real outputs, producing insight the team acts on. Days 60–90 — Scale & Systemize: Make the framework reusable across capabilities and wire it into the model/data/eval loop. What You Bring 5+ years in product evaluation, UX research, model testing, or similar structured qualitative assessment. Master's or higher in Cognitive Science, HCI, Design Research, Psychology, Media Studies, or a related field. Deep familiarity with creative workflows for generative models (animation, filmmaking, digital art, VFX). Systems thinking: you can define abstract qualities like believability or scene coherence in clear evaluative terms. Excellent written communication and the ability to synthesize nuanced judgment into actionable insight. Comfort working across engineers, researchers, and creatives. Nice to Have Background in motion, visual effects, or storytelling pipelines. Experience evaluating AI-generated media (video, images, 3D). Prior work building internal tools for qualitative data collection or scoring. Familiarity with prompt engineering and reference-based inputs. About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
Every tech & IT company hiring across India — with AI match scores — on one live map.
Open the map →