Benchmark

PythonSaga

Benchmark for evaluating code-generating large language models.

Theme
Evaluation, Reliability & Trustworthy AI
Type
Benchmark
Funding
Unfunded

Paper and data

PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs

Yadav, Beniwal, Singh. Findings of EMNLP 2024.

PaperarXivGitHubHugging Face

Theme →Datasets and models →Publications →