Pac-Bench is a new benchmark designed to evaluate the code generation capabilities of AI models. Specifically, it focuses on how effectively models can produce a complete and functional Pac-Man game from a minimal input.
The benchmark provides each AI model with a single prompt: "Create a Pac-Man game in a single HTML page." This strict methodology means models are not allowed any follow-up prompts, clarifications, or opportunities to correct their initial output. This approach aims to test the models' ability to generate complex code accurately in a single attempt.
This benchmark offers insights into the one-shot code generation performance of AI models. Understanding how well models can interpret a high-level request and translate it into a complete, working application without iterative refinement is relevant for developing more autonomous code generation tools.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Pac-Bench is a new benchmark that tests how well AI models can create a functional Pac-Man game from a single prompt. Each model is given one attempt without follow-up prompts or corrections, assessing their one-shot code generation capabilities.