Researchers introduced ClinLens, a benchmark of 200 tasks testing AI agents on complex longitudinal clinical data analysis across electronic health records, medical imaging, and notes from the MIMIC dataset. Current leading AI models achieve only 56.3% accuracy on clinical tasks despite producing executable code 100% of the time, exposing a critical gap between code that runs and code that produces clinically correct analyses.
Why it matters: As healthcare systems increasingly adopt AI for data analysis, this benchmark reveals that execution success doesn't guarantee clinical correctness—a distinction critical for practitioners evaluating AI coding agents for patient-facing applications.