Automatic sentence simplification system for Arabic Script Punjabi

T Tayyaba Shehzad S Sadaf Abdul Rauf S Saleha Nazeer A Ali Daud H Hussain Dawood

Abstract

In the domain of language simplification, creating aligned monolingual parallel datasets tailored to specific linguistic dialects is a significant endeavor. This pursuit, introduces the pioneering Punjabi Simplification (PUSIM) corpus, which focuses on the Shahmukhi dialect. Shahmukhi, one of the two prominent dialects of Punjabi, serves as the foundation for this corpus development. This study employs a hybrid approach that ensures the comprehensive assessment of simplification outcomes. The detailed process of simplification underwent thorough examination, aiming to transform complex sentences into simpler ones by enhancing writing clarity and vocabulary. To quantify the quality and readability of simplified texts, automated readability assessments were conducted using well-established text readability metrics, a significant SARI score of 45.3, attested to the high quality of simplification approach. The unique aspect of this work lies in its focus on the Shahmukhi dialect, addressing a linguistic facet that had previously received limited attention in natural language processing. It is anticipated that this dataset will pave the way for further exploration and research, offering novel possibilities for leveraging automated simplification techniques in the realm of Shahmukhi Punjabi language processing.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 6
Published June 11, 2026
Pages e0344915
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (5)

T

Tayyaba Shehzad

S

Sadaf Abdul Rauf

S

Saleha Nazeer

A

Ali Daud

H

Hussain Dawood