Wiese, T. (2026). Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge. PLoS ONE. https://doi.org/10.1371/journal.pone.0339920