I advise health systems, foundations, nonprofits, and health AI companies on the design, performance, and safety evaluation of AI systems in healthcare, including predictive models and LLM-based applications. This includes support with evaluation design, red teaming and safety/bias assessment, human-AI interaction review, and guidance on AI governance before and after deployment. Past and ongoing partners include health systems (Emory Healthcare) and nonprofit organizations (e.g. Myna Mahila Foundation, ARMMAN, UN Global Pulse).

Get in touch

Areas of support

Evaluation design

Metrics, human evaluation protocols, and LLM-as-a-judge setups matched to real-world use. See our rapid review of 116 LLM evaluation metrics in healthcare (EFMI STC 2026).

Benchmarks

Test sets co-designed with patients, health workers, and clinicians to reflect real settings.

Safety and bias

Red teaming and bias assessments for sensitive health topics and diverse populations.

Human-AI interaction

Interfaces and workflows that fit how people use, question, and rely on AI.

Participatory design

Sessions that bring communities, care workers, and administrators into shaping AI tools from the start.

Governance

Processes for reviewing and monitoring AI tools before and after deployment.

Ways to work together

  • Advisory calls and roles

    Focused calls on an evaluation plan or deployment decision, or ongoing roles as a board or technical advisor.

  • Scoped projects

    Time-bound work on one AI system, such as an evaluation, safety review, or workshop, with clear recommendations.

  • Research partnerships

    Longer-term collaborations through the CARE Lab, from co-designed studies to evaluations in real-world deployment.

I offer flexible arrangements for nonprofits and organizations in low- and middle-income countries.

Get in touch