Research and career
Lakkaraju's doctoral research focused on developing and evaluating interpretable, transparent, and fair predictive models which can assist human decision makers (e.g., doctors, judges) in domains such as healthcare, criminal justice, and education. As part of her doctoral thesis, she developed algorithms for automatically constructing interpretable rules for classification and other complex decisions which involve trade-offs. Lakkaraju and her co-authors also highlighted the challenges associated with evaluating predictive models in settings with missing counterfactuals and unmeasured confounders, and developed new computational frameworks for addressing these challenges. She co-authored a study which demonstrated that when machine learning models are used to assist in making bail decisions, they can help reduce crime rates by up to 24.8% without exacerbating racial disparities.
Lakkaraju joined Harvard University as a postdoctoral researcher in 2018, and then became an assistant professor at the Harvard Business School and the Department of Computer Science at Harvard University in 2020. Over the past few years, she has done pioneering work in the area of explainable machine learning. She initiated the study of adaptive and interactive post hoc explanations which can be used to explain the behavior of complex machine learning models in a manner that is tailored to user preferences. She and her collaborators also made one of the first attempts at identifying and formalizing the vulnerabilities of popular post hoc explanation methods. They demonstrated how adversaries can game popular explanation methods, and elicit explanations that hide undesirable biases (e.g., racial or gender biases) of the underlying models. Lakkaraju also co-authored a study which demonstrated that domain experts may not always interpret post hoc explanations correctly, and that adversaries could exploit post hoc explanations to manipulate experts into trusting and deploying biased models.
She also worked on improving the reliability of explanation methods. She and her collaborators developed novel theory and methods to analyze and improve the robustness of different classes of post hoc explanation methods by proposing a unified theoretical framework and establishing the first known connections between explainability and adversarial training. Lakkaraju has also made important research contributions to the field of algorithmic recourse. She and her co-authors developed one of the first methods which allows decision makers to vet predictive models thoroughly to ensure that the recourse provided is meaningful and non-discriminatory. Her research has also highlighted critical flaws in several popular approaches in the literature of algorithmic recourse.