Vector-based code retrieval is a critical building block in code assistants and agents. However, many people complained about the lack of diverse, high-quality evaluation datasets for it. We surveyed existing ones and proposed some methods to build better ones. 🧵🧵

We identify 3 main issues with existing datasets: noisy labels, simplism, and data contamination risks (see above). We propose two methods to create new datasets, repurposing QA datasets and leveraging the issue/ticket record, and use them to build many new datasets below.

We evaluated various embedding models, @OpenAI , @awscloud CodeSage, CodeRankEmbed, @JinaAI_ v2 code, along with the @Voyage AI’s newly released voyage-code-3 (https://t.co/0lL99O7eSR) on these datasets:

Please check out a longer version in our blog post on code evaluation: https://t.co/WylTF7m5Kz
If you choose a code model from public benchmarks, this names the flaws in those datasets and gives two reproducible ways to build sets from your own QA data and issue history.
Checking sign-in…
Loading comments…