
An open training recipe for retrieval models lets teams build code and long- search without depending on closed data.
“State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap.”
“They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class.”
“Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts.”
“This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe.”
Checking sign-in…
Loading comments…