Evaluating the cultural alignment of multilingual LLMs in typical Japanese workplace scenarios
Abstract
While current evaluations of LLM cultural alignment predominantly rely on static benchmarks in Western contexts, their ability to navigate generative, high-context socio-pragmatic demands in non-Western environments remains critically underexplored. This study investigates how multilingual LLMs adapt to the Japanese workplace—a stringent stress-test environment characterized by strong high-context communication norms and rigid honorific conventions—using Hofstede’s six cultural dimensions as a heuristic framework. We evaluated five state-of-the-art LLMs (LLM-jp, Phi, Llama, Qwen, and GLM) through a large-scale crowdsourced human evaluation. Based on 1,718 valid evaluator sessions, native Japanese raters assessed model outputs to generate a holistic Japanese Workplace Cultural Alignment Score (JWCAS). To dissect the underlying communicative strategies, we paired this with a three-layer diagnostic sub-score analysis (Linguistic Form, Socio-Cultural Values, and Social Action). Our results reveal that leading multilingual models (Phi and GLM) achieved overall JWCAS scores comparable to, or significantly higher than, the native Japanese model (LLM-jp). Crucially, our sub-score analysis demonstrates that holistic evaluation metrics can obscure deep pragmatic deficits: while LLM-jp overfits to surface-level linguistic politeness (Layer 1), it shows critical weaknesses in socio-cultural values (Layer 2) and context-aware social strategies (Layer 3). In contrast, leading multilingual models demonstrate balanced competence across all layers. These findings suggest that true cultural competence requires moving beyond native linguistic mastery, highlighting the necessity of multi-dimensional diagnostic frameworks for cross-cultural AI alignment.
Article Details
Authors (6)
Zhiwei Gao
Nobuyuki Shimizu
Sumio Fujita
Shaowen Peng
Shoko Wakamiya
Eiji Aramaki