Thomas Wiese. "Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge." PLoS ONE, 2026. doi:10.1371/journal.pone.0339920