Researchers benchmarked five watermarking schemes across 18 language models and found that watermarks—used to trace AI-generated content—substantially degrade performance on medical tasks, causing lexical errors, hallucinated terminology, and missed diagnostic findings. The study introduces the first human-expert-validated evaluation framework for detecting watermark-induced failures in clinical text, revealing that standard benchmarks systematically mask these degradations.
Why it matters: As LLMs enter clinical workflows, this research demonstrates that current watermarking approaches may introduce dangerous failure modes in medical AI without detection—making domain-specific safety evaluation mandatory before deployment in healthcare.