Sign inLaunch

FrontierHarness Eval vs PacBench

Both are Coding & Dev Tools and Research & Data tools. FrontierHarness Eval is open source; PacBench is free.

FrontierHarness EvalPacBench
TaglineSame model, 12 harness configurations, up to 17.5x cost differenceHow well can models one-shot a Pac-Man game?
PricingOpen SourceFree
CategoriesCoding & Dev Tools, Research & DataCoding & Dev Tools, Research & Data
Built forDevelopers, ResearchersDevelopers
PlatformsWeb, CLIWeb
Tech stack——
Alternative to——
AI Launch upvotes▲ 0▲ 0
Hacker NewsY▲ 82 on HNY▲ 78 on HN
Launched2026-10-022026-09-23

Overview

Who it's for

FrontierHarness Eval

Engineers choosing or building a coding-agent harness.

PacBench

Developers and AI enthusiasts comparing how well models build a working game in one shot.

Problem

FrontierHarness Eval

It's hard to tell how much of an agent's cost and quality comes from the harness rather than the model.

PacBench

Abstract benchmark scores don't show what a model can actually build from a single prompt.

Solution

FrontierHarness Eval

A controlled evaluation where every run uses the same model, hardware and fresh checkpoint, so only the harness changes.

PacBench

A gallery of one-shot Pac-Man builds with the cost, time and token counts of each run.

What makes it unique

FrontierHarness Eval

All 360 trials start from the same fresh checkpoint restore to avoid warm-cache bias, and you can run it on your own harness.

PacBench

Every entry answers the same one-line prompt, so results are directly comparable and playable.