01The Role
Tenera helps teams make product decisions using simulated users. Understanding when those simulations can be trusted raises difficult questions about measurement, evidence, and generalization.
As a Research Scientist in Evaluations, you will own an ambitious, open-ended research agenda with substantial uncertainty. The questions, methods, and path to a useful result will not arrive fully defined. You will help define them, develop original approaches, and test your ideas with scientific rigor.
This role requires exceptional creativity and execution in equal measure. You will have significant autonomy and sustained time for independent research, with responsibility for turning uncertain ideas into concrete progress. You should be comfortable building experiments yourself, seeking critical feedback, and pursuing difficult questions even when early attempts fail.
You will have a dedicated research budget to collect data and build benchmarks, with the autonomy and responsibility to decide how those resources can best advance your research.
02What You'll Do
- Define and pursue an ambitious research agenda in simulation evaluation. Decide which questions matter, challenge existing assumptions, and develop original approaches when established methods are insufficient.
- Turn loosely defined questions into concrete research plans. Set milestones, prioritize experiments, and revise your direction as evidence emerges.
- Design and implement evaluation methods grounded in observed behavior. Build the datasets, experimental tools, and analyses needed to test your ideas end to end.
- Build and operate data-labeling pipelines, including annotation guidelines, annotator calibration, quality checks, and adjudication. Establish traceable ground truth and document uncertainty or disagreement in the labels.
- Design controlled comparisons, holdout datasets, and replication studies. Account for sampling bias, data leakage, confounding, and uncertainty before drawing conclusions.
- Study validity, reliability, and generalization across populations and contexts. Seek out counterexamples, investigate failures, and distinguish a promising result from a defensible conclusion.
- Own execution from the first hypothesis through reproducible results and clear scientific writing. Seek targeted criticism from teammates while remaining responsible for research direction and progress.
03What We're Looking For
- A track record of original empirical or methodological research in statistics, machine learning, behavioral science, psychometrics, computational social science, or a related field. Demonstrate your contribution through publications, research reports, or substantial independent projects.
- Exceptional creativity. You can identify questions others have overlooked, connect ideas across disciplines, and devise useful methods when familiar approaches do not work.
- Exceptional execution. You turn ambitious ideas into working experiments, resolve practical obstacles, and carry difficult projects through to concrete, verifiable results.
- Comfort with substantial uncertainty and sustained independent work. You can make progress without a detailed brief, tolerate inconclusive experiments, and change course without losing sight of the research goal.
- Strong foundations in statistical inference and experimental design. You can reason about measurement validity, statistical power, uncertainty, selection bias, and the limits of causal claims.
- Prior experience building evaluation benchmarks end to end is required, from defining the task and sampling data to establishing ground truth, designing scoring methods, and validating the benchmark.
- Prior experience building and operating data-labeling pipelines is required. You have written annotation guidelines, calibrated annotators, measured agreement, resolved ambiguous labels, and implemented quality controls and dataset versioning.
- Strong Python and data analysis skills. You can implement research methods, work with real-world datasets, and build reproducible experiments that others can inspect.
- Clear scientific writing and intellectual honesty. You seek criticism, revise your conclusions, and report null results and limitations with the same care as positive findings.
- Humble. You are quick to learn from customers, teammates, and evidence instead of defending your first answer.
- Low ego. You care more about the best idea winning than being personally right.
04How to Apply
Tenera is an in-person company, working 5 days a week in our SF office. Open to relocation. Meaningful equity for the right person.