Flores-Herr, Nicolas
YOU?
Author Swipe
View article: Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs Open
We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe’s linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-Englis…
View article: Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models Open
High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly r…
View article: Towards End-to-End Model-Agnostic Explanations for RAG Systems
Towards End-to-End Model-Agnostic Explanations for RAG Systems Open
Retrieval Augmented Generation (RAG) systems, despite their growing popularity for enhancing model response reliability, often struggle with trustworthiness and explainability. In this work, we present a novel, holistic, model-agnostic, po…
View article: Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis
Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis Open
Chunking is a crucial preprocessing step in retrieval-augmented generation (RAG) systems, significantly impacting retrieval effectiveness across diverse datasets. In this study, we systematically evaluate fixed-size chunking strategies and…
View article: Fact Finder - Enhancing Domain Expertise of Large Language Models by Incorporating Knowledge Graphs
Fact Finder - Enhancing Domain Expertise of Large Language Models by Incorporating Knowledge Graphs Open
Recent advancements in Large Language Models (LLMs) have showcased their proficiency in answering natural language queries. However, their effectiveness is hindered by limited domain-specific knowledge, raising concerns about the reliabili…
View article: Towards Multilingual LLM Evaluation for European Languages
Towards Multilingual LLM Evaluation for European Languages Open
The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European languages remains …