for teams
View source Published October 5, 2026

Introducing bon travail

Agents find broken code and pay humans to fix it.

1. Introduction

Today, we are making agents pay humans.

bon travail is an open-source system in which an agent finds broken tests and a human is paid to fix them. The agent, Aeon, watches a team’s CI. When the same test fails twice, it works out why. The engineer decides whether the fix stays in the team or goes to a human. The repository’s own tests decide whether the fix works. The human is paid in USDC from an escrow on Arc testnet, with a sealed receipt.

We ran this end to end on October 4, 2026. A regression was traced, handed out, fixed, merged, verified on the default branch, and paid onchain. A second task was refunded when its deadline passed.

2. How it works

The GitHub App reads workflow runs, logs, contents, and pull requests on the repositories a team connects. It never writes. A job and step that fail twice in a row become a finding, with GitHub’s log and the commits between the last green run and the first red one.

Aeon then reproduces the failure in its own runner, narrows it to the first bad commit, and proposes how a fix should be judged. The engineer reads that and keeps the work internal, dismisses it, or hands it out with a reward and a deadline. Once handed out, a human claims it by opening a pull request. The tests run. The reward is released or refunded. Aeon keeps watching for the failure to come back.

3. Who decides what

Agents describe. People and tests decide. Every decision that moves money belongs to someone other than the agent.

DecisionWho decides
Which repositories are watchedThe engineer
Whether a finding leaves the teamThe engineer
Reward, deadline, scope, and protected pathsThe engineer, frozen at approval
Who may take the workThe engineer: named GitHub logins, or anyone on GitHub
Whether a fix worksThe repository’s own GitHub Actions
Releasing or refunding moneyThe escrow contract, only after the verdict

Aeon can trigger a sweep, read findings, and attach an investigation. Anything it sends about rewards, people, or approval is dropped. An agent that could choose recipients or amounts would turn a prompt injection in a log or pull request into a payment. Keeping it descriptive means the worst a hostile repository can do is mislead an investigation that an engineer then reads.

4. Verification

The verdict comes from GitHub’s records. The pull request has to target the right repository and branch, come from the claimant, and leave .github/, the watched workflow, and every protected path untouched. By default the fix must also be merged, and the run that counts is the default branch’s own run on the merge commit.

Nothing a contributor, an engineer’s browser, or Aeon says can change that verdict.

5. Payments and escrow

Rewards sit in the ProofworkEscrow contract on Arc testnet at 0xe816…4dd8. A task is funded when the engineer approves, released to the claimant’s wallet when the verdict passes, and refunded to the funder when the deadline passes. Each task can settle once.

Before launch, Circle’s Arc Studio reviewed the deployed contract from its bytecode, and each finding was checked against the source and the live chain:

FindingVerdict
Owner and operator are the same keyTrue. Acceptable on testnet; split before mainnet.
The reward cap is fixed at 1 USDCTrue. It is immutable; production needs a redeploy.
A refund may go to the callerFalse. Refunds pay the stored funder.
A task could pay twiceNot possible. Release and refund both require a funded task and close it first.
Ownership could be renouncedFalse. It reverts, checked live.

The server also caps the total held in escrow at once, so a stolen operator key has a bounded blast radius.

6. First live run

On October 4, 2026 a regression was pushed to a public sandbox, usdc-sdk-examples. Aeon traced it to commit 4721bb1 with high confidence. The work was handed out for 0.25 USDC, fixed in pull request #3, merged, verified by the repository’s tests on main, and paid onchain in block 65461543.

The full record is on its receipt. The same loop also exercised a refund, when a deadline passed with no fix.

7. Why this matters

The usual story is that agents replace the engineer. The systems worth building do the opposite. They watch, explain, and pay a person for the work they cannot close. The person keeps the decision.

That is how the people building these systems describe the goal. Altman has said the aim is tools that elevate people, not entities that replace them1. Suleyman describes superintelligence that always works in service of people2. Nadella calls AI a scaffolding for human potential rather than a substitute3. Karpathy’s version is a computable metric on one side and a human who still has to check on the other4. Allaire argues that agents and onchain settlement are one economy, in which agents pay for outcomes5, and Circle opened Arc’s public mainnet in September with USDC as its gas6.

A failing test is a small instance of that economy.

The work is already large, and it is unowned. At Google, about 1.5% of test runs were flaky, and about 84% of the transitions from passing to failing involved a flaky test rather than a real bug. Each one still needed a person to look7. Poor software quality cost the United States an estimated $2.41 trillion in 20228. Developers report about 17 hours a week on maintenance, including about four on bad code, instead of new work9.

Agents are good at finding this and explaining it. Humans are good at fixing it. Tests are good at judging it. bon travail connects the three, and the escrow settles the bill only after the repository’s own run passes. The agent does not choose the recipient, the amount, or the verdict. It cannot talk a test into passing.

8. Build it with us

bon travail is open source and live today. The next stage needs more than code, and we are looking for people to build it with.

  • Teams. If your team has tests that keep breaking, run the loop on your repositories with us and shape what it becomes.
  • Funding. Partners and grants to grow the reward pool, so the humans who fix real code are paid in real USDC.
  • Builders. Engineers and agent builders who want agents that hire people rather than replace them.

Write to us at hello@bontravail.xyz, or open an issue on GitHub.

9. References

  1. S. Altman, post on X, May 1, 2026, on building tools that elevate people rather than replace them. Source ↩
  2. M. Suleyman, “Towards Humanist Superintelligence,” Microsoft AI, November 2025. Source ↩
  3. S. Nadella, “Looking Ahead to 2026,” sn scratchpad, December 2025. Source ↩
  4. A. Karpathy, on autonomous research and removing yourself as the bottleneck, March 2026. Source ↩
  5. J. Allaire, “The Agentic Economy: The Convergence of Intelligence and the Economy,” July 2026. Source ↩
  6. Circle launches Arc public mainnet, with USDC as gas, September 16, 2026. Source ↩
  7. J. Micco, “Flaky Tests at Google and How We Mitigate Them,” Google Testing Blog, May 2016. Source ↩
  8. CISQ, “The Cost of Poor Software Quality in the US: A 2022 Report,” December 2022. Source ↩
  9. Stripe, “The Developer Coefficient,” 2018. Source ↩