Thomas Wiese. "Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge." PLoS ONE (2026). https://doi.org/10.1371/journal.pone.0339920