Usage-Surplus OSS Program

Surplus usage should become public goods.

Expiring coding-agent capacity becomes tested maintenance for active open-source projects—not another unused meter rolling back to zero.

OSS · FIELD UNIT 01Usage-Surplus OSS Program raccoon field operative carrying code patches and a wrench
Special Agent, Upstream Affairs

Origin

Five-day meters.
Zero appetite for waste.

Many subscription coding plans operate on recurring five-day usage windows. We often reach the end of a window with unused capacity spread across Claude Max, Codex and ChatGPT Pro, SuperGrok, and whichever capable coding-agent plan is currently in rotation.

The program began with a simple refusal to waste that surplus. The first Claude missions went into open-source evaluation and agent infrastructure: reproduce a real failure, read the contribution rules, make the smallest responsible repair, and leave behind evidence a maintainer can inspect.

Missions can be a sharply bounded bug fix, a new feature, benchmark-adapter work, difficult compatibility repair, tests, documentation, or the deeply unglamorous maintenance that keeps widely used infrastructure alive.

Not this program

PR-volume farming · drive-by patch spam · bounty labor · abandoned-repository code dumping · generated code without accountable review

Operating model

How a deployment works

Selection is constrained by maintainer activity, reproducibility, public value, and our ability to verify the work honestly.

01

Weekly scan

We look for substantive open bugs, missing compatibility work, and maintenance gaps in public projects we use or understand well enough to test responsibly.

02

Maintainer-activity check

A project qualifies only when maintainers are demonstrably present: reviewing contributions, giving direction, and merging acceptable work within a reasonable period.

03

Reproduce before editing

The assigned agent reads the repository rules, reproduces the issue on the supported stack, and records the failure mode before proposing a patch.

04

Implement and verify natively

The patch follows the project's own architecture and runs its native tests, lint, typing, packaging, or end-to-end validation—not a substitute invented for the contribution.

05

Human review gate

Kirsten reviews the mission, diff, evidence, and public claims. No submission is marked maintainer-ready until a human accepts responsibility for it.

06

Disclose and follow through

Commits and PRs disclose agent involvement. We answer review, rebase when needed, rerun checks, and stay with the contribution after the exciting part is over.

Mission lanes

Where surplus capacity goes

Public-interest engineering with a bias toward evaluation validity, agent reliability, defensive research, and technical agency.

01

Open-source benchmarks

Correctness, reproducibility, packaging, environment, and verifier maintenance for benchmarks that people still rely on—including older versions that remain mainstream long after a newer paper or release appears.

Age does not make a benchmark irrelevant when its scores still shape decisions.
02

Evaluation ecosystems and adapters

Public benchmark adapters for frameworks such as NeMo Gym, plus careful maintenance when the benchmark science changes. A 2024 benchmark may later gain a corrected dataset, revised exclusions, new task splits, a successor protocol, or a different LLM-as-a-Judge rubric and judge model.

Responsible adapters add explicit, provenance-preserving support for the new publication without silently changing the meaning of historical scores.
03

Agent harnesses and frameworks

OpenHands, Pi and kimi-pi, DSH, Harbor, and adjacent runtimes: unattended execution, tool behavior, lifecycle evidence, portability, failure semantics, and integration with evaluation systems.

The model is not the whole agent. Harness behavior belongs in the evidence.
04

ML infrastructure and serving machinery

SGLang, vLLM, model-serving layers, sandboxes, and controlled security-research runtimes that expose gradients, logits, activations, or other attack-relevant signals safely enough to study adversarial behavior.

This lane builds defensive and safety-research infrastructure—not operational attack enablement.
05

Access and technical self-determination

Nonprofits and community groups teaching girls, women, and trans people to write agents, deploy and fine-tune models, and own the infrastructure rather than merely consume it.

In the spirit of Grace Hopper, with the profanity restored: Just because we have always done it this way does not mean we have to fucking keep doing it this way.

Current fieldwork

Public work, linked directly.

These are live public records—not a claim that every draft has merged or that fork review is the same as upstream acceptance.

Public integration fork

NeMo Gym program work

The following packages are public PRs to reinainblood/Gym, the integration and review fork. They are not represented here as upstream NVIDIA PRs.

Provenance and accountability

No mystery authorship.
No invisible submissions.

Kirsten chooses the missions and remains accountable for work submitted under her account. Agent authorship is disclosed in commits and PRs. Public descriptions identify what was reproduced, what changed, and which project-native checks were run.

A draft can be useful evidence of active work; it is not maintainer approval. Human review is required before ready-for-review status. Review feedback, rebases, reruns, and follow-up fixes are part of the mission rather than an optional epilogue.

Exact test receiptsDisclosed agent authorshipHuman submission gateMaintainer follow-through

Request a deployment

Have an active project and a real blocked issue?

Send the repository URL, issue or bug link, reproduction evidence if available, why the work is currently blocked, evidence that maintainers are active, and any contribution or security constraints.

The program is non-paid, capacity-dependent, and public-interest oriented. It is not a commercial service, a support contract, or a guarantee that a contribution will be selected, accepted, or merged.

Email a deployment request

Drafted by Special Agent Claude of the Usage-Surplus OSS Program.

No tokens left behind.