Research Publications

Scientific contributions to the field of machine translation and natural language processing

Highlighted Publications

ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT

ACL 2025

Mikołaj Pokrywka, Wojciech Kusa, Mieszko Rutkowski, Mikołaj Koszowski

View Publication

Laniqo at WMT25 Terminology Translation Task: A Multi-Objective Reranking Strategy for Terminology-Aware Translation via Pareto-Optimal Decoding

WMT 2025

Kamil Guttmann, Adrian Charkiewicz, Zofia Rostek, Mikołaj Pokrywka, Artur Nowakowski

View Publication

Adam Mickiewicz University at WMT 2022: NER-Assisted and Quality-Aware Neural Machine Translation

WMT 2022

Artur Nowakowski, Gabriela Pałka, Kamil Guttmann, Mikołaj Pokrywka

View Publication

All Publications

2026
2025
2024
2023
2022
2021

2026 Publications

Findings of the WMT26 Terminology Translation Task: The Hard Part is Finding the Terms, Not Using Them

WMT 2026, October 2026

Adrian Charkiewicz, Pinzhen Chen, Thierry Etchegoyhen, Harritxu Gete, Kamil Guttmann, Xu Huang, David Ponce, Artur Nowakowski, Frédéric Odermatt, Arturo Oncevay, Dawei Zhu, Vilém Zouhar, Kirill Semenov

The WMT26 Terminology Translation Task aims to evaluate machine translation in high-stakes, term-heavy domains (technology, finance, medicine). This year, we focus solely on document-level translation and run two tasks: (1) MT with explicit document-level dictionaries, (2) MT with bitext samples that contain the specific terms. Participants are presented with the texts in three translation directions, two of which feature mid-to-low-resourced morphologically rich languages: Spanish→Basque, English→Polish, and Traditional Chinese→English. This year, the main metrics were multiple variants of overall translation quality and terminology success rate; in line with previous shared tasks, we also compared systems with no terminology, proper dictionaries, and random dictionaries to causally analyze terminology utility. 17 teams participated in our task, submitting 21 systems to Track 1 and 18 systems to Track 2. The results show that document-level translation with explicit dictionaries is close to saturation, and the best systems nearly reach the reference texts. In contrast, for translation with sample bitexts, the spread of the systems is bigger, and the best scores are lower, highlighting the need to concentrate on terminology extraction rather than its use from an explicit source. We also evaluate the grammaticality of the generated texts in Basque and Polish and observe slight trends toward using dictionary forms of the terms and assigning the most frequent grammemes.

Laniqo at WMT26 Multilingual Instruction Shared Task (MIST): Full-Parameter Knowledge Distillation with Task-Conditional Pivot Decoding Under a 10B-Parameter Constraint

WMT 2026, October 2026

Adrian Charkiewicz, Artur Nowakowski

  • 1st place · Open-Ended Generation, human evaluation — top-ranked system in the official human-evaluation ranking Source
  • 1st place · Cross-lingual Summarization track — highest Language-Gated ROUGE-L across 24 languages Source

We describe our submission to the WMT26 Multilingual Instruction Shared Task (MIST). Our system is google/gemma-4-E4B-it (8B total / 4.5B effective parameters), fully fine-tuned via knowledge distillation from a larger teacher model (Gemma-4-31B-it) on teacher-generated responses spanning three sub-tasks and 24 languages. We find that (i) a two-judge quality-agreement filter on the distillation data provides no measurable benefit over using the teacher’s outputs unfiltered, at both LoRA and full-parameter training scale, and (ii) a task-conditional decoding strategy (using two-turn English drafting for open-ended generation and summarization, versus single-turn direct decoding for context-grounded question answering) significantly improves quality where applied and is harmless where it is not. Our final submission combines full-parameter distillation with this task-conditional decoding scheme.

Laniqo at WMT26 Video Subtitle Translation Shared Task: Multi-Objective Fusion of Pareto-Optimal Candidates

WMT 2026, October 2026

Kamil Guttmann, Artur Nowakowski

  • 2nd place · Video Subtitle Translation Shared Task — second overall of five teams; 2nd in Indonesian, Malay, Thai and Traditional Chinese Source

This paper describes Laniqo’s submission to the WMT26 Video Subtitle Translation Shared Task, translating subtitles from Simplified Chinese into English, Thai, Indonesian, Malay, and Traditional Chinese under a 20B-parameter, open-license model constraint. Our system optimizes translation quality and subtitle compliance entirely at inference time, without fine-tuning, by combining hedged multi-prompt candidate generation, language-identification and compliance pruning, multi-objective Pareto reranking, reasoning-based candidate fusion, and a fallback/compression step. We screened four open-weight models and compared the two strongest, Qwen3.5-9B and Gemma-4-12B-it, across the full pipeline on the complete test corpus. Gemma-4-12B-it reached significantly higher compliance rates in every target language, thus it was chosen as our final system.

Multi-Candidate Synthesis in Multimodal Machine Translation

FedCSIS 2026, August 2026

Jan Kolwicz, Szymon Niewiadomski, Kamil Guttmann, Michał Ciesiółka, Krzysztof Jassem

Multimodal Machine Translation (MMT) typically relies on passing text and corresponding images to a model to resolve ambiguities. However, the direct presence of an image can occasionally mislead the model, causing regressions compared to text-only translation. To address this, we propose a 5-step pipeline for evaluating and refining MMT. Instead of single-pass translations, our approach generates four diverse candidate translations leveraging different subsets of multimodal context (text-only, text+image, text+caption, text+image+caption). A final synthesis step acts as a judge and editor to produce the optimal translation. Evaluated on the ConECT and Multi30k datasets using large language models (Qwen3.6-35B and Gemma 4 26B), our synthesis pipeline consistently outperforms text-only and standard multimodal approaches. We observe up to a +1.34 improvement in COMET scores.

CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs

EAMT 2026, June 2026

Kamil Guttmann, Zofia Fraś, Artur Nowakowski, Krzysztof Jassem

Current state-of-the-art Quality Estimation (QE) in machine translation relies on massive, proprietary LLMs, raising data privacy concerns. We demonstrate that smaller, open-source LLMs (<30B parameters) are a viable, cost-effective and privacy-preserving alternative. Using a single-pass prompting strategy, our models simultaneously generate quality scores, MQM error annotations, suggested error corrections, and full post-editions. Our analysis shows these models achieve highly competitive system-level correlations with human judgments that outperform traditional neural metrics, fine-tuned models, and human inter-annotator agreement, effectively approximating the capabilities of much larger proprietary LLMs.

ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

EAMT 2026, June 2026

Michał Ciesiółka, Dawid Wiśniewski, Adrian Charkiewicz, Kamil Guttmann

We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure structural diversity in the dataset, we employ K-Medoids sampling over 45 geometric features, capturing complex elements like nested tables and formulas to focus only on visually diverse PDF documents. Our evaluation reveals that current MT systems struggle with spatial grounding and geometric synchronization, often losing the link between text and its visual context. ForMaT provides a benchmark for developing layout-aware translation models that integrate visual and textual context for high-fidelity document reconstruction.

2025 Publications

Laniqo at WMT25 Terminology Translation Task: A Multi-Objective Reranking Strategy for Terminology-Aware Translation via Pareto-Optimal Decoding

WMT 2025, November 2025

Kamil Guttmann, Adrian Charkiewicz, Zofia Rostek, Mikołaj Pokrywka, Artur Nowakowski

  • 1st place · Terminology Translation Task, Track 1 — rank 1 on the quality–terminology Pareto front Source

This paper describes the Laniqo system submitted to the WMT25 Terminology Translation Task. Our approach uses a Large Language Model fine-tuned on parallel data augmented with source-side terminology constraints. To select the final translation from a set of generated candidates, we introduce Pareto-Optimal Decoding – a multi-objective reranking strategy. This method balances translation quality with term accuracy by leveraging several quality estimation metrics alongside Term Success Rate (TSR). Our system achieves TSR greater than 0.99 across all language pairs on the Shared Task testset, demonstrating the effectiveness of the proposed approach.

Laniqo at WMT25 General Translation Task: Self-Improved and Retrieval-Augmented Translation

WMT 2025, November 2025

Kamil Guttmann, Zofia Rostek, Adrian Charkiewicz, Antoni Solarski, Mikołaj Pokrywka, Artur Nowakowski

This work describes Laniqo’s submission to the constrained track of the WMT25 General MT Task. We participated in 11 translation directions. Our approach combines several techniques: fine-tuning the EuroLLM-9B-Instruct model using Contrastive Preference Optimization on a synthetic dataset, applying RetrievalAugmented Translation with human-translated data, implementing Quality-Aware Decoding, and performing postprocessing of translations with a rule-based algorithm. We analyze the contribution of each method and report improvements at every stage of our pipeline.

ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT

ACL 2025, July 2025

Mikołaj Pokrywka, Wojciech Kusa, Mieszko Rutkowski, Mikołaj Koszowski

Neural Machine Translation (NMT) has improved translation by using Transformer-based models, but it still struggles with word ambiguity and context. This problem is especially important in domain-specific applications, which often have problems with unclear sentences or poor data quality. Our research explores how adding information to models can improve translations in the context of e-commerce data. To this end we create ConECT — a new Czech-to-Polish e-commerce product translation dataset coupled with images and product metadata consisting of 11,400 sentence pairs. We then investigate and compare different methods that are applicable to context-aware translation. We test a vision-language model (VLM), finding that visual context aids translation quality. Additionally, we explore the incorporation of contextual information into text-to-text models, such as the product's category path or image descriptions. The results of our study demonstrate that the incorporation of contextual information leads to an improvement in the quality of machine translation. We make the new dataset publicly available.

Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation — a Multilingual Perspective

MT Summit 2025, June 2025

Dawid Wiśniewski, Mikołaj Pokrywka, Zofia Rostek

Current machine translation models provide us with high-quality outputs in most scenarios. However, they still face some specific problems, such as detecting which entities should not be changed during translation. In this paper, we explore the abilities of popular NMT models, including models from the OPUS project, Google Translate, MADLAD, and EuroLLM, to preserve entities such as URL addresses, IBAN numbers, or emails when producing translations between four languages: English, German, Polish, and Ukrainian. We investigate the quality of popular NMT models in terms of accuracy, discuss errors made by the models, and examine the reasons for errors. Our analysis highlights specific categories, such as emojis, that pose significant challenges for many models considered. In addition to the analysis, we propose a new multilingual synthetic dataset of 36,000 sentences that can help assess the quality of entity transfer across nine categories and four aforementioned languages.

Exploring the Feasibility of Multilingual Grammatical Error Correction with a Single LLM up to 9B parameters: A Comparative Study of 17 Models

MT Summit 2025, June 2025

Dawid Wiśniewski, Antoni Solarski, Artur Nowakowski

Recent language models can successfully solve various language-related tasks, and many understand inputs stated in different languages. In this paper, we explore the performance of 17 popular models used to correct grammatical issues in texts stated in English, German, Italian, and Swedish when using a single model to correct texts in all those languages. We analyze the outputs generated by these models, focusing on decreasing the number of grammatical errors while keeping the changes small. The conclusions drawn help us understand what problems occur among those models and which models can be recommended for multilingual grammatical error correction tasks. We list six models that improve grammatical correctness in all four languages and show that Gemma 9B is currently the best performing one for the languages considered.

2024 Publications

FAME-MT Dataset: Formality Awareness Made Easy for Machine Translation Purposes

EAMT 2024, June 2024

Dawid Wiśniewski, Zofia Rostek, Artur Nowakowski

People use language for various purposes. Apart from sharing information, individuals may use it to express emotions or to show respect for another person. In this paper, we focus on the formality level of machine-generated translations and present FAME-MT – a dataset consisting of 11.2 million translations between 15 European source languages and 8 European target languages classified to formal and informal classes according to target sentence formality. This dataset can be used to fine-tune machine translation models to ensure a given formality level for 8 European target languages considered. We describe the dataset creation procedure, the analysis of the dataset's quality showing that FAME-MT is a reliable source of language register information, and we construct a publicly available proof-of-concept machine translation model that uses the dataset to steer the formality level of the translation. Currently, it is the largest dataset of formality annotations, with examples expressed in 112 European language pairs. The dataset is made available online.

Chasing COMET: Leveraging Minimum Bayes Risk Decoding for Self-Improving Machine Translation

EAMT 2024, June 2024

Kamil Guttmann, Mikołaj Pokrywka, Adrian Charkiewicz, Artur Nowakowski

This paper explores Minimum Bayes Risk (MBR) decoding for self-improvement in machine translation (MT), particularly for domain adaptation and low-resource languages. We implement the self-improvement process by fine-tuning the model on its MBR-decoded forward translations. By employing COMET as the MBR utility metric, we aim to achieve the reranking of translations that better aligns with human preferences. The paper explores the iterative application of this approach and the potential need for language-specific MBR utility metrics. The results demonstrate significant enhancements in translation quality for all examined language pairs, including successful application to domain-adapted models and generalisation to low-resource settings. This highlights the potential of COMET-guided MBR for efficient MT self-improvement in various scenarios.

2023 Publications

Exploring the Use of Foundation Models for Named Entity Recognition and Lemmatization Tasks in Slavic Languages

EACL 2023, May 2023

Gabriela Pałka, Artur Nowakowski

  • 1st place · Name normalization, Slav-NER 2023 — best normalization F1 in Czech, Polish and Russian Source

This paper describes Adam Mickiewicz University's (AMU) solution for the 4th Shared Task on SlavNER. The task involves the identification, categorization, and lemmatization of named entities in Slavic languages. Our approach involved exploring the use of foundation models for these tasks. In particular, we used models based on the popular BERT and T5 model architectures. Additionally, we used external datasets to further improve the quality of our models. Our solution obtained promising results, achieving high metrics scores in both tasks. We describe our approach and the results of our experiments in detail, showing that the method is effective for NER and lemmatization in Slavic languages. Additionally, our models for lemmatization will be available at: https://huggingface.co/amu-cai.

2022 Publications

Adam Mickiewicz University at WMT 2022: NER-Assisted and Quality-Aware Neural Machine Translation

WMT 2022, September 2022

Artur Nowakowski, Gabriela Pałka, Kamil Guttmann, Mikołaj Pokrywka

  • 1st place · General MT Task, Czech↔Ukrainian Source

This paper presents Adam Mickiewicz University's (AMU) submissions to the constrained track of the WMT 2022 General MT Task. We participated in the Ukrainian ↔ Czech translation directions. The systems are a weighted ensemble of four models based on the Transformer (big) architecture. The models use source factors to utilize the information about named entities present in the input. Each of the models in the ensemble was trained using only the data provided by the shared task organizers. A noisy back-translation technique was used to augment the training corpora. One of the models in the ensemble is a document-level model, trained on parallel and synthetic longer sequences. During the sentence-level decoding process, the ensemble generated the n-best list. The n-best list was merged with the n-best list generated by a single document-level model which translated multiple sentences at a time. Finally, existing quality estimation models and minimum Bayes risk decoding were used to rerank the n-best list so that the best hypothesis was chosen according to the COMET evaluation metric. According to the automatic evaluation results, our systems rank first in both translation directions.

POLENG MT: An Adaptive MT Platform

EAMT 2022, June 2022

Artur Nowakowski, Krzysztof Jassem, Maciej Lison, Kamil Guttmann, Mikołaj Pokrywka

We introduce POLENG MT, an MT platform that may be used as a cloud web application or as an on-site solution. The platform is capable of providing accurate document translation, including the transfer of document formatting between the input document and the output document. The main feature of the on-site version is dedicated customer adaptation, which consists of training on specialized texts and applying forced terminology translation according to the user's needs.

nEYron: Implementation and Deployment of an MT System for a Large Audit & Consulting Corporation

EAMT 2022, June 2022

Artur Nowakowski, Krzysztof Jassem, Maciej Lison, Rafał Jaworski, Tomasz Dwojak, Karolina Wiater, Olga Posesor

This paper reports on the implementation and deployment of an MT system in the Polish branch of EY Global Limited. The system supports standard CAT and MT functionalities such as translation memory fuzzy search, document translation and post-editing, and meets less common, customer-specific expectations. The deployment began in August 2018 with a Proof of Concept, and ended with the signing of the Final Version acceptance certificate in October 2021. We present the challenges that were faced during the deployment, particularly in relation to the security check and installation processes in the production environment.

2021 Publications

Adam Mickiewicz University's English-Hausa Submissions to the WMT 2021 News Translation Task

WMT 2021, November 2021

Artur Nowakowski, Tomasz Dwojak

This paper presents the Adam Mickiewicz University's (AMU) submissions to the WMT 2021 News Translation Task. The submissions focus on the English↔Hausa translation directions, which is a low-resource translation scenario between distant languages. Our approach involves thorough data cleaning, transfer learning using a high-resource language pair, iterative training, and utilization of monolingual data via back-translation. We experiment with NMT and PB-SMT approaches alike, using the base Transformer architecture for all of the NMT models while utilizing PB-SMT systems as comparable baseline solutions.

Approaching English-Polish Machine Translation Quality Assessment with Neural-based Methods

PolEval 2021, October 2021

Artur Nowakowski

This paper presents our contribution to the PolEval 2021 Task 2: Evaluation of translation quality assessment metrics. We describe experiments with pre-trained language models and state-of-the-art frameworks for translation quality assessment in both nonblind and blind versions of the task. Our solutions ranked second in the nonblind version and third in the blind version.

Neural Machine Translation with Inflected Lexicon

MT Summit 2021, August 2021

Artur Nowakowski, Krzysztof Jassem

The paper presents experiments in neural machine translation with lexical constraints into a morphologically rich language. In particular and we introduce a method and based on constrained decoding and which handles the inflected forms of lexical entries and does not require any modification to the training data or model architecture. To evaluate its effectiveness and we carry out experiments in two different scenarios: general and domain-specific. We compare our method with baseline translation and i.e. translation without lexical constraints and in terms of translation speed and translation quality. To evaluate how well the method handles the constraints and we propose new evaluation metrics which take into account the presence and placement and duplication and inflectional correctness of lexical terms in the output sentence.

A Neural Translator Designed to Protect the Eastern Border of the European Union

MT Summit 2021, August 2021

Artur Nowakowski, Krzysztof Jassem

This paper reports on a translation engine designed for the needs of the Polish State Border Guard. The engine is a component of the AI Searcher system, whose aim is to search for Internet texts, written in Polish, Russian, Ukrainian or Belarusian, which may lead to criminal acts at the eastern border of the European Union. The system is intended for Polish users, and the translation engine should serve to assist understanding of non-Polish documents. The engine was trained on general-domain texts. The adaptation for the criminal domain consisted in the appropriate translation of criminal terms and proper names, such as forenames, surnames and geographical objects. The translation process needs to take into account the rich inflection found in all of the languages of interest. To this end, a method based on constrained decoding that incorporates an inflected lexicon into a neural translation process was applied in the engine.