Rendered at 07:00:16 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
yathern 7 hours ago [-]
The winning model (Cloud Agent) is Promptless, and in very small grey text, you might be able to spot a "Built by Promptless" on the page if you look closely.
While I'm sure this benchmark is one worth optimizing against, and some real thought was put into it - it also seems possible that the rubric was specifically designed to be one that helps promote its creator. If one runs a contest, it's a dangerous situation to also have an entrant. Either the contest can be favored for the entrant, or the entrant can have special training privileges. To put it another way, it can't be assumed you'll have a fair trial when your judge is also your uncle.
Not saying that this has happened here, but third party, unaffiliated benchmarks are always preferred for this reason.
frances-liu 6 hours ago [-]
Appreciate the feedback! It's totally valid.
1. I just made the "by Promptless" darker on the website so that it's more obvious :).
2. Re: rubrics design, agreed that there's definitely conflict of interest, and we were actively trying to figure out how to navigate that ourselves. (In the paper we excluded cloud agents, including Promptless, for example.)
We originally wanted to create DoGBench because we were seeing teams proclaim that documentation is a solved problem by AI, and in some cases even lay off entire technical writing teams (most infamously at Snowflake). Meanwhile, in our own efforts to build a docs agent, we were seeing all kinds of cases where AI was struggling. So we wanted to quantify just how much human judgement is still valuable in this field.
A few notes on the methodology that might be worth mentioning:
- The tasks are all from open-source repos; there are no synthetic tasks
- The rubric generation process was validated with real technical writers who are open-source maintainers (at Helm, PostHog, and Mautic).
- We put the source code and many of the tasks/rubrics/trajectories online, and anyone can submit a new or updated agent for judging
We actually considered putting task and rubric assembly entirely in the hands of independent 3rd-party and just submit Promptless results like how anyone would. But we didn't have enough resources to assemble this independent committee when we started this project. Now that DoGBench is gaining some traction in the documentation domain, maybe we'll do that for v2!
frances-liu 8 hours ago [-]
Hi HN, I'm the first author, ask me anything. I've done AI research before but this is our first time building an agent benchmark so lots of lessons learned. If you're curious about AI docs agents or curious about the benchmark itself, ask away!
While I'm sure this benchmark is one worth optimizing against, and some real thought was put into it - it also seems possible that the rubric was specifically designed to be one that helps promote its creator. If one runs a contest, it's a dangerous situation to also have an entrant. Either the contest can be favored for the entrant, or the entrant can have special training privileges. To put it another way, it can't be assumed you'll have a fair trial when your judge is also your uncle.
Not saying that this has happened here, but third party, unaffiliated benchmarks are always preferred for this reason.
1. I just made the "by Promptless" darker on the website so that it's more obvious :).
2. Re: rubrics design, agreed that there's definitely conflict of interest, and we were actively trying to figure out how to navigate that ourselves. (In the paper we excluded cloud agents, including Promptless, for example.)
We originally wanted to create DoGBench because we were seeing teams proclaim that documentation is a solved problem by AI, and in some cases even lay off entire technical writing teams (most infamously at Snowflake). Meanwhile, in our own efforts to build a docs agent, we were seeing all kinds of cases where AI was struggling. So we wanted to quantify just how much human judgement is still valuable in this field.
A few notes on the methodology that might be worth mentioning:
- The tasks are all from open-source repos; there are no synthetic tasks
- The rubric generation process was validated with real technical writers who are open-source maintainers (at Helm, PostHog, and Mautic).
- We put the source code and many of the tasks/rubrics/trajectories online, and anyone can submit a new or updated agent for judging
We actually considered putting task and rubric assembly entirely in the hands of independent 3rd-party and just submit Promptless results like how anyone would. But we didn't have enough resources to assemble this independent committee when we started this project. Now that DoGBench is gaining some traction in the documentation domain, maybe we'll do that for v2!