← All stories
● Covered by 1 source · 1 reportLow impact1 neutral

Pac-Bench evaluates AI models' ability to generate Pac-Man game from single prompt

🔄 Updated 3d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Pac-Bench evaluates AI models on Pac-Man game generation.
  • Models receive a single prompt: "Create a Pac-Man game in a single HTML page".
  • No follow-up prompts or fixes are allowed.
  • The benchmark assesses one-shot code generation.

Introducing Pac-Bench

Pac-Bench is a new benchmark designed to evaluate the code generation capabilities of AI models. Specifically, it focuses on how effectively models can produce a complete and functional Pac-Man game from a minimal input.

Evaluation Methodology

The benchmark provides each AI model with a single prompt: "Create a Pac-Man game in a single HTML page." This strict methodology means models are not allowed any follow-up prompts, clarifications, or opportunities to correct their initial output. This approach aims to test the models' ability to generate complex code accurately in a single attempt.

Significance for AI Development

This benchmark offers insights into the one-shot code generation performance of AI models. Understanding how well models can interpret a high-level request and translate it into a complete, working application without iterative refinement is relevant for developing more autonomous code generation tools.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Pac-Bench is a new benchmark that tests how well AI models can create a functional Pac-Man game from a single prompt. Each model is given one attempt without follow-up prompts or corrections, assessing their one-shot code generation capabilities.