I advise health systems, foundations, nonprofits, and health AI companies on designing and evaluating AI systems for safe, effective use in real-world care, with a special focus on large language model (LLM) applications.
Benchmarks alone rarely show how a system will behave with the people it is meant to serve. For companies, I help rigorously test whether products work safely and equitably across diverse patients, clinicians, and settings, before and after launch. For foundations and nonprofits, I help ensure AI programs work for the communities they serve, including in low-resource settings.
My approach draws on my research on AI evaluation as a sociotechnical practice and on prior and ongoing partnerships with organizations such as Emory Healthcare, ARMMAN, and Myna Mahila Foundation.
Get in touchWhat I help with
Evaluation design
Metrics, human evaluation protocols, and LLM-as-a-judge setups matched to real-world use.
Benchmarks
Test sets co-designed with patients, health workers, and clinicians to reflect real settings.
Safety and bias
Red teaming and bias assessments for sensitive health topics and diverse populations.
Human-AI interaction
Interfaces and workflows that fit how people use, question, and rely on AI.
Participatory design
Sessions that bring communities, care workers, and administrators into shaping AI tools from the start.
Governance
Processes for reviewing and monitoring AI tools before and after deployment.
Ways to work together
-
Advisory calls and roles
Focused calls on an evaluation plan or deployment decision, or ongoing roles as a board or technical advisor.
-
Scoped projects
Time-bound work on one AI system, such as an evaluation, safety review, or workshop, with clear recommendations.
-
Research partnerships
Longer-term collaborations through the CARE Lab, from co-designed studies to evaluations in real-world deployment.
I offer flexible arrangements for nonprofits and organizations in low- and middle-income countries.