As AI is continually deployed in healthcare settings, all stakeholders (model builders, adopters of AI solutions) need to understand how models are performing in real-world settings.
But most evaluations used today are built on synthetic data or human generated data. Although useful in certain contexts, these data often miss the reality of how healthcare is performed in real life. This causes a clinical readiness gap, which is the delta between how a model performs on evaluations versus how it performs when deployed inside a real clinical workflow. A recent review published in JAMA found that only about 5% of recent LLM evaluations in medicine were conducted on actual EHR data. Because real-world data, such as electronic health records, carry the ground truth of what happened in a clinical context, models can be tested against actual clinical workflows.
Protege provides access to multi-modal, real-world healthcare data that is representative of clinical and administrative tasks in the real world, creating benchmarks that more accurately reflect how health AI will continue to be used in real life settings.