Research

Vision

Progress in AI is measured through experiments, and experiments are only as trustworthy as the methods used to run and interpret them. My research asks a simple question that turns out to be hard: when an NLP paper, a benchmark, or a deployed system claims that one model is better than another, or that a model can replace a human, how do we know that the claim is true?

My answer draws on statistics, experimental design, and machine learning. I develop methods that let researchers and practitioners test their claims properly: significance tests suited to NLP data, protocols for comparing models across many datasets and many prompts, measures of annotation quality, and statistical procedures for deciding when an LLM can stand in for human annotators or judges. I then put these methods to work on real problems, often with collaborators from psychology, education, law, and the life sciences.

Looking ahead, I want to set standards for hybrid workflows in which humans and LLMs work together, so that the gains in speed and cost do not come at the price of data quality or place a heavier cognitive load on the people involved. I am also interested in modeling and evaluating LLM personas, with the goal of building personalized models whose behavior is validated rather than assumed. The broader conviction behind this agenda is that rapid advances in AI must be matched by careful protocols, and eventually regulation, that make its deployment scientifically rigorous, ethical, and cost-effective across society.

Research topics

Statistically sound evaluation in NLP

Much of my early work addressed how NLP results are compared and reported. This includes guidance on choosing the right significance test for a given evaluation measure, replicability analysis for conclusions drawn from multiple datasets, and a criterion for comparing deep neural models that accounts for the variance across random seeds and hyperparameters. These ideas are collected in a book on statistical significance testing for NLP. With colleagues at the University of Pennsylvania, I extended this line of work to the meta-evaluation of automatic metrics, showing how much uncertainty surrounds correlations between metrics and human judgments.

  • The Hitchhiker's Guide to Testing Statistical Significance in NLP (ACL 2018)
  • Replicability Analysis for NLP: Testing Significance with Multiple Datasets (TACL 2017)
  • Deep Dominance: How to Properly Compare Deep Neural Models (ACL 2019)
  • Statistical Significance Testing for Natural Language Processing (Morgan & Claypool, 2020)
  • A Statistical Analysis of Summarization Evaluation Metrics Using Resampling Methods (TACL 2021)

Evaluating large language models

LLMs are sensitive to how they are prompted, and evaluations built on a single prompt can rank models in ways that do not hold up. I work on evaluation protocols that account for this variability, on the limits of reference-free evaluation of generated text, and on benchmarks for complex, goal-driven settings such as web agents.

Data annotation and LLM-as-a-judge

LLMs are increasingly used to label data and to judge the outputs of other models. Whether they can replace human annotators is a statistical question, and I develop tests that answer it with explicit guarantees. Related work develops frameworks for measuring the quality and consistency of human annotators, including in subjective tasks where there is no single correct label. This direction is the focus of my Israel Science Foundation grant, Data Annotation and Model Evaluation in the Era of Large Language Models.

  • The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs (ACL 2025)

LLM personas and role-playing agents

Simulated users and role-playing agents are useful for research and for applications such as education and mental health, but only if their behavior is consistent and believable. I study how to build personas grounded in psychology and how to measure whether an agent actually maintains the persona it was given, including the naturalness of its dialogue.

Language technology for people and society

I collaborate with researchers in other disciplines to apply NLP, and to evaluate it rigorously, in settings where the stakes are human. Current projects include Hebrew text simplification for accessibility, LLM-based adaptive surveys that measure extreme psychological responses, detection of behavioral and emotional themes in text, LLM-supported computer science education, and the analysis of public attitudes expressed on social media.

The complete list is on the publications page.