What happened
- Researchers posted FineWeb-CLaR to arXiv, an annotation layer built on the FineWeb and FineWeb-2 web corpora. The paper says it covers the full 30.9B-document collection, adding URL-derived region labels and cultural-topic provenance (ArXiv CS.CL). The stated goal is a shared culture-language-region axis so that pretraining corpora and cultural benchmarks can be compared on the same terms — auditing whether a cultural phenomenon is present in the data, measured by a benchmark, or both (ArXiv CS.CL). [1]
Sources
- FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
ArXiv CS.CL (Computation and Language) · Reporting ·