Weekly scan
We look for substantive open bugs, missing compatibility work, and maintenance gaps in public projects we use or understand well enough to test responsibly.
Usage-Surplus OSS Program
Expiring coding-agent capacity becomes tested maintenance for active open-source projects—not another unused meter rolling back to zero.

Origin
Many subscription coding plans operate on recurring five-day usage windows. We often reach the end of a window with unused capacity spread across Claude Max, Codex and ChatGPT Pro, SuperGrok, and whichever capable coding-agent plan is currently in rotation.
The program began with a simple refusal to waste that surplus. The first Claude missions went into open-source evaluation and agent infrastructure: reproduce a real failure, read the contribution rules, make the smallest responsible repair, and leave behind evidence a maintainer can inspect.
Missions can be a sharply bounded bug fix, a new feature, benchmark-adapter work, difficult compatibility repair, tests, documentation, or the deeply unglamorous maintenance that keeps widely used infrastructure alive.
PR-volume farming · drive-by patch spam · bounty labor · abandoned-repository code dumping · generated code without accountable review
Operating model
Selection is constrained by maintainer activity, reproducibility, public value, and our ability to verify the work honestly.
We look for substantive open bugs, missing compatibility work, and maintenance gaps in public projects we use or understand well enough to test responsibly.
A project qualifies only when maintainers are demonstrably present: reviewing contributions, giving direction, and merging acceptable work within a reasonable period.
The assigned agent reads the repository rules, reproduces the issue on the supported stack, and records the failure mode before proposing a patch.
The patch follows the project's own architecture and runs its native tests, lint, typing, packaging, or end-to-end validation—not a substitute invented for the contribution.
Kirsten reviews the mission, diff, evidence, and public claims. No submission is marked maintainer-ready until a human accepts responsibility for it.
Commits and PRs disclose agent involvement. We answer review, rebase when needed, rerun checks, and stay with the contribution after the exciting part is over.
Mission lanes
Public-interest engineering with a bias toward evaluation validity, agent reliability, defensive research, and technical agency.
Correctness, reproducibility, packaging, environment, and verifier maintenance for benchmarks that people still rely on—including older versions that remain mainstream long after a newer paper or release appears.
Age does not make a benchmark irrelevant when its scores still shape decisions.Public benchmark adapters for frameworks such as NeMo Gym, plus careful maintenance when the benchmark science changes. A 2024 benchmark may later gain a corrected dataset, revised exclusions, new task splits, a successor protocol, or a different LLM-as-a-Judge rubric and judge model.
Responsible adapters add explicit, provenance-preserving support for the new publication without silently changing the meaning of historical scores.OpenHands, Pi and kimi-pi, DSH, Harbor, and adjacent runtimes: unattended execution, tool behavior, lifecycle evidence, portability, failure semantics, and integration with evaluation systems.
The model is not the whole agent. Harness behavior belongs in the evidence.SGLang, vLLM, model-serving layers, sandboxes, and controlled security-research runtimes that expose gradients, logits, activations, or other attack-relevant signals safely enough to study adversarial behavior.
This lane builds defensive and safety-research infrastructure—not operational attack enablement.Nonprofits and community groups teaching girls, women, and trans people to write agents, deploy and fine-tune models, and own the infrastructure rather than merely consume it.
In the spirit of Grace Hopper, with the profanity restored: Just because we have always done it this way does not mean we have to fucking keep doing it this way.Current fieldwork
These are live public records—not a claim that every draft has merged or that fork review is the same as upstream acceptance.
Unattended execution contract for benchmark and CI harnesses
View public PR ↗02Public OpenHands CLI adapter with ATIF trajectories and Modal V2 validation
View public PR ↗03Correct round-tripping for versioned package datasets
View public PR ↗04AssistantBench scorer crash on unhashable set-literal answers
View public PR ↗05Per-sample canary accounting for prompt-injection exfiltration
View public PR ↗06Reject blank keyword and regex checks that otherwise pass everything
View public PR ↗07Preserve dataset configuration through kwargs constructor chains
View public PR ↗08Remove a broken conda shell hook from 300 task images
View public PR ↗09Prevent task images from exposing future Git history
View public PR ↗Public integration fork
The following packages are public PRs to reinainblood/Gym, the integration and review fork. They are not represented here as upstream NVIDIA PRs.
Provenance and accountability
Kirsten chooses the missions and remains accountable for work submitted under her account. Agent authorship is disclosed in commits and PRs. Public descriptions identify what was reproduced, what changed, and which project-native checks were run.
A draft can be useful evidence of active work; it is not maintainer approval. Human review is required before ready-for-review status. Review feedback, rebases, reruns, and follow-up fixes are part of the mission rather than an optional epilogue.
Request a deployment
Send the repository URL, issue or bug link, reproduction evidence if available, why the work is currently blocked, evidence that maintainers are active, and any contribution or security constraints.
The program is non-paid, capacity-dependent, and public-interest oriented. It is not a commercial service, a support contract, or a guarantee that a contribution will be selected, accepted, or merged.
Email a deployment requestDrafted by Special Agent Claude of the Usage-Surplus OSS Program.
No tokens left behind.