For a full list, see Google Scholar.
Some harmful model capabilities are difficult or unsafe to evaluate by generating outputs, especially when producing the content itself is legally or ethically constrained, as with child sexual abuse material (CSAM). We develop methods for detecting harmful model specialization from a model's internal representations, enabling evaluation without generation.
AI systems increasingly personalize their responses using information about users and their prior interactions, but this context can also amplify problematic forms of agreement and mirroring. We show that interaction context often increases sycophancy, highlighting how context-free evaluations can underestimate behaviors that raise concerns about user autonomy and psychological harm in real-world interactions.
Differential privacy makes it possible to learn from sensitive data while providing formal guarantees about individual privacy, but these guarantees can come at a substantial cost to statistical accuracy. We develop Iterative Hessian Mixing, a Gaussian-sketching approach to private linear regression that achieves near-optimal accuracy and improves over state-of-the-art methods.
What equal opportunity requires cannot always be determined by evaluating individual algorithmic decisions in isolation; it depends on how systems collectively structure access to opportunity. We develop algorithmic pluralism as a normative and evaluative framework, arguing that systems should preserve multiple viable pathways to opportunity and be evaluated by the severity and legitimacy of the bottlenecks they create.