Leadership wants to declare a redesign a success three weeks after launch based on a bump in daily active users. What's the most rigorous stance?
- A. Agree, since DAU is up and that's what matters
- B. Caution that early novelty effects and confounds may inflate DAU, and propose tracking task-success and retention over a longer window against a baseline ✓
- C. Insist on waiting a full year before saying anything
- D. Pick whichever metric currently looks best and report that
Correct answer: B. Attributing impact rigorously requires controlling for novelty and confounds and measuring outcome metrics over a meaningful horizon.
Your research strongly recommends against a feature the CEO is championing. How do you exercise influence without authority?
- A. Comply silently since the CEO will win regardless
- B. Present the evidence and its business risk clearly, propose a low-cost experiment to test the CEO's hypothesis, and let results guide the call ✓
- C. Circulate the findings widely to build opposition to the CEO
- D. Soften the findings until they no longer conflict with the CEO's view
Correct answer: B. Influencing without authority means making the risk legible and offering a fair, low-cost test rather than either capitulating or fighting.
An A/B test shows a statistically significant 0.3% lift on a huge sample. The PM wants to ship. What's the most senior read?
- A. Ship immediately; significance means it's real and beneficial
- B. Weigh whether the effect size is practically meaningful and worth the added complexity and long-term cost, not just statistically significant ✓
- C. Reject it because the lift is small
- D. Re-run the test indefinitely until the number grows
Correct answer: B. Senior judgment distinguishes statistical significance from practical significance and weighs it against costs and complexity.
Two valid studies you ran point to opposite recommendations for the same decision. How do you handle it?
- A. Go with the study whose result the team prefers
- B. Examine method differences, context, and what each study is actually measuring to reconcile them into a nuanced recommendation ✓
- C. Average the two recommendations into a compromise design
- D. Throw out both and declare research inconclusive
Correct answer: B. Conflicting evidence calls for investigating why they differ and synthesizing a context-aware recommendation, not picking the convenient one.
You must choose a north-star metric for a new product with your leadership team. What's the strongest principle to advocate?
- A. Pick the metric that grows fastest so the team looks good
- B. Choose a metric that captures genuine user value delivered and is hard to game by degrading the experience ✓
- C. Choose whatever competitors use as their north star
- D. Use raw signups because it's simple to explain
Correct answer: B. A good north-star metric reflects real user value and resists being juiced at the user's expense.
A long-running satisfaction metric is trending down but you can't tell why from the numbers alone. As research lead, what's the best move?
- A. Report the decline and let the team speculate on causes
- B. Launch targeted qualitative research to diagnose drivers and segment where the decline concentrates ✓
- C. Wait another quarter to see if it recovers on its own
- D. Change to a different metric that looks healthier
Correct answer: B. Diagnosing a metric decline requires qualitative and segmentation work to find causal drivers, not more speculation.
A powerful stakeholder repeatedly cherry-picks single quotes from your studies to justify their agenda. How do you address it long-term?
- A. Stop sharing raw quotes so they can't misuse them
- B. Establish a shared practice of presenting findings with prevalence and context, and reinforce the weight of evidence behind each claim ✓
- C. Publicly call out the stakeholder for misusing research
- D. Only present findings that the stakeholder can't twist
Correct answer: B. Building norms around weighting findings by evidence and context structurally reduces cherry-picking better than withholding or confrontation.
You're deciding between a longitudinal diary study and a one-time usability test for understanding how a habit-forming feature is adopted. What drives the choice?
- A. Pick the usability test because it's faster and cheaper
- B. Choose based on the question: adoption over time and changing behavior needs longitudinal data a single session can't capture ✓
- C. Always prefer the diary study since it collects more data
- D. Let the budget alone decide regardless of the question
Correct answer: B. Method selection follows the phenomenon: behavior that unfolds over time requires a longitudinal design, not a snapshot.
Your study sample skews toward one region, but the product is global. Leadership wants to generalize the findings worldwide. What's your responsibility?
- A. Let them generalize since the findings are strong
- B. Clearly bound the claims to the population studied and advocate for follow-up research in other markets before global decisions ✓
- C. Refuse to share the findings at all
- D. Quietly reword the findings to sound global
Correct answer: B. Rigorous researchers scope claims to the studied population and flag where generalization is unsupported.
A team wants to run continuous A/B tests on a vulnerable user population (e.g., people in financial distress). What ethical stance do you take?
- A. Proceed since experimentation is standard practice
- B. Weigh the potential for harm, ensure safeguards and informed consideration of impact, and limit experiments that could exploit vulnerability ✓
- C. Ban all testing with these users permanently
- D. Test freely as long as conversion improves
Correct answer: B. Research ethics require heightened care and harm-avoidance when experimenting on vulnerable populations, not blanket permission or prohibition.
Your qualitative insight contradicts a well-established quantitative model the data-science team trusts. How do you proceed?
- A. Concede since quantitative data outranks qualitative
- B. Explore whether the two are measuring different things, and collaborate to reconcile the behavioral 'why' with the statistical 'what' ✓
- C. Insist your qualitative finding is correct and theirs is flawed
- D. Drop your finding to avoid conflict with data science
Correct answer: B. Qualitative and quantitative methods often illuminate different facets; reconciling them collaboratively yields the fullest truth.
You're asked to demonstrate research ROI to justify the team's budget. What's the most defensible approach?
- A. Count the number of studies run and hours logged
- B. Trace specific product decisions that changed because of research and the value or risk avoided as a result ✓
- C. Show how satisfied stakeholders are with your reports
- D. Emphasize how much data you collected overall
Correct answer: B. Research value is best demonstrated through decisions influenced and risk avoided, not activity volume.
A design tests well in the lab but you suspect it may fail in messy real-world contexts. What's the strongest next step?
- A. Trust the lab results since they were controlled
- B. Validate in a more realistic setting or via field/analytics data before committing, since ecological validity is uncertain ✓
- C. Ship it and treat production as the test
- D. Redesign it based on your suspicion without further evidence
Correct answer: B. Lab success doesn't guarantee real-world performance; testing ecological validity de-risks the gap before commitment.
Two segments of users need opposite things from the same feature. How do you guide the product decision?
- A. Design for the larger segment and ignore the smaller
- B. Quantify the size and value of each segment, surface the trade-off explicitly, and help the team choose or differentiate deliberately ✓
- C. Try to satisfy both fully in one design
- D. Let the loudest segment's feedback decide
Correct answer: B. Making the segment trade-off explicit with sizing lets the team make a deliberate strategic choice rather than an accidental one.
Analytics show a feature is barely used. The PM concludes users don't want it and wants to kill it. What's the more rigorous interpretation?
- A. Agree; low usage clearly means low demand
- B. Investigate whether low usage reflects low value or poor discoverability, onboarding, or fit before concluding ✓
- C. Keep the feature regardless since someone built it
- D. Assume the analytics tracking is simply broken
Correct answer: B. Low usage is ambiguous — it can mean low value or poor execution — so it must be diagnosed before a kill decision.
You want research embedded earlier in the product process, but teams only call you to 'validate' finished designs. How do you shift this?
- A. Keep doing validation studies and hope it changes
- B. Demonstrate the value of early discovery on one project and build repeatable rituals that pull research into problem definition ✓
- C. Refuse all late-stage validation requests until they change
- D. Complain to leadership that teams misuse research
Correct answer: B. Changing when research is used is best achieved by proving upstream value and institutionalizing it, not by refusing work.
A survey and a behavioral log disagree about how often users perform an action. Which do you generally trust more, and why?
- A. The survey, because users know their own behavior best
- B. The behavioral log for frequency, since self-report is prone to recall and social-desirability bias — while using the survey to understand perception ✓
- C. Whichever gives the more convenient answer
- D. Neither; discard both as unreliable
Correct answer: B. For actual frequency, behavioral data beats self-report, though surveys remain valuable for perceptions and reasons.
Leadership pushes for a big generative research initiative, but the immediate roadmap decisions need answers in two weeks. How do you balance strategy and delivery?
- A. Drop the tactical needs to invest fully in the strategic study
- B. Sequence the work: deliver fast targeted studies for imminent decisions while scoping the deeper generative research in parallel ✓
- C. Refuse the generative work as a distraction
- D. Do the generative study slowly and let roadmap decisions wait
Correct answer: B. Senior researchers serve both horizons by triaging urgent decisions while investing in strategic understanding in parallel.
You uncover that a growth tactic driving good metrics relies on a mildly deceptive pattern. The team is thrilled with the numbers. What do you do?
- A. Stay quiet since the metrics are strong and it's not your call
- B. Surface the ethical and long-term trust risk with evidence, and advocate for a sustainable alternative ✓
- C. Immediately report the team to leadership for wrongdoing
- D. Accept it because growth is the priority this quarter
Correct answer: B. Researchers are stewards of the user's interest and must raise trust and ethics risks, framed constructively with alternatives.
Your team must decide whether to invest in a redesign. You can run one rigorous study or three quick-and-dirty ones. How do you decide?
- A. Always choose the single rigorous study for credibility
- B. Match the design to the decision's stakes and reversibility: high-stakes irreversible calls justify rigor, while several cheap probes may better de-risk an exploratory bet ✓
- C. Always choose the three quick studies to cover more ground
- D. Let whoever is loudest in the room decide the format
Correct answer: B. Calibrating research depth to the decision's stakes and reversibility is a hallmark of senior research judgment.