← Trở về danh mục
HỒ SƠ TÀI LIỆU · #14.364
Chỉ metadata Tài liệu khác EN cc-by-4.0

A Rule-Based Approach for Aligning Japanese-Spanish Sentences from A Comparable Corpora

Năm xuất bản2023
Nguồn học thuậtZenodo
Định danhDOI 10.5121/ijnlc.2012.1301
TRÍCH DẪN ĐỀ XUẤT

International Journal on Natural Language Computing (IJNLC) (2023). A Rule-Based Approach for Aligning Japanese-Spanish Sentences from A Comparable Corpora. Zenodo. https://doi.org/10.5121/ijnlc.2012.1301

01

Tóm tắt

The performance of a Statistical Machine Translation System (SMT) system is proportionally directed to the quality and length of the parallel corpus it uses. However for some pair of languages there is a considerable lack of them. The long term goal is to construct a Japanese-Spanish parallel corpus to be used for SMT, whereas, there are a lack of useful Japanese-Spanish parallel Corpus. To address this problem, In this study we proposed a method for extracting Japanese-Spanish Parallel Sentences from Wikipedia using POS tagging and Rule-Based approach. The main focus of this approach is the syntactic features of both languages. Human evaluation was performed over a sample and shows promising results, in comparison with the baseline.

02

Thông tin thư mục