What the preprint found
A new arXiv preprint (2609.26942v1, ArXiv CS.CL) benchmarks five open-weight LLMs on kinship terms in Hindi, Tamil and Korean, testing generation against a matched four-option selection baseline. On identical relation-language cells, GPT OSS-120B picked the correct term in 90.67% of 75 valid cells but produced an accepted term in 36.00% of the corresponding attempts, while Llama 3.3-70B showed the same gap (77.92% vs 24.24%). The authors say the gap is best read as an evaluation-format effect, since the four-option setup shows the candidates and requires no script production, not as proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy ranged from GLM-5.1 at 72.29% to Llama-3.3-70B at 24.24%. The paternal-lineage advantage was language-specific — large in Hindi, weak or reversed in Korean — with Tamil shared-term pairs used as a control for measurement variation. The authors conclude culturally specific kinship generation stays hard even when the relationship is spelled out, and call for generation-based evaluation alongside multiple-choice tests. [1]
Sources
- Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
ArXiv CS.CL (Computation and Language) · Reporting ·