Sections

Search

ArXiv paper measures how much benchmark contamination actually inflates scoresGoogle's Gemini 3.8 TTS models are cheap and handle multi-voice dialogue, per Simon WillisonUniDataAgent: China Unicom's Ontology-Grounded Enterprise Q&A Agent Cuts Report Time from Days to MinutesOpenAI says Harvey uses GPT-6 Astra to produce more structured legal draftsArXiv paper proposes auditable LLM labeling for classroom talk
All stories

Models·

FineWeb gets culture, language and region labels to audit what pretraining data misses

arXiv paper releases FineWeb-CLaR, annotating 30.9B web documents and 277 cultural NLP benchmarks on a shared culture-language-region axis.

What happened

- Researchers posted FineWeb-CLaR to arXiv, an annotation layer built on the FineWeb and FineWeb-2 web corpora. The paper says it covers the full 30.9B-document collection, adding URL-derived region labels and cultural-topic provenance (ArXiv CS.CL). The stated goal is a shared culture-language-region axis so that pretraining corpora and cultural benchmarks can be compared on the same terms — auditing whether a cultural phenomenon is present in the data, measured by a benchmark, or both (ArXiv CS.CL). [1]

Sources

  1. ArXiv CS.CL (Computation and Language) · Reporting ·
    FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing