I just came across ProgramBench, a new benchmark from the creators of SWEBench. Definitely not perfect, but a good addition to this post. Here's a nice thread with what is this all about and common criticism to it: https://x.com/deedydas/status/2051684179084284409
A brand new (and interesting) benchmark just dropped, DeepSWE. Worth including it to the list from this post: https://deepswe.datacurve.ai/blog
I just came across ProgramBench, a new benchmark from the creators of SWEBench. Definitely not perfect, but a good addition to this post. Here's a nice thread with what is this all about and common criticism to it: https://x.com/deedydas/status/2051684179084284409
Somewhat tangential, but I came across this interesting post from Anthropic about agent evaluation that can be a great complement to this post: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents